VLDB 2026 Research / reviewers in the wild / expert
Dani Lischinski
dblp:29/19
· DBLP profile ↗
127ranked-venue papers
4as first author
36since 2021 · last 2026
0000-0002-6191-0361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 37 · 21 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Iconix: Controlling Semantics and Style in Progressive Icon Grids GenerationabstractVisual communication often needs stylistically consistent icons that span concrete and abstract meanings, for use in diverse contexts. We present Iconix, a human-AI co-creative system that organizes icon generation along two axes: semantic richness (what is depicted) and visual complexity (how much detail). Given a user-specified concept, Iconix constructs a semantic scaffold of related analytical perspectives and employs chained, image-conditioned generation to produce a coherent style of exemplars. Each exemplar is then automatically distilled into a progressive sequence, from detailed and elaborate to abstract and simple. The resulting two-dimensional grid exposes a navigable space, helping designers reason jointly about figurative content and visual abstraction. A within-subjects study (N = 32) found that compared to a baseline workflow, participants produced icon grids more creatively, reported lower workload, and explored a coherent range of design variations. We discuss implications for human-machine co-creative approaches that couple semantic scaffolding with progressive simplification to support visual abstraction. Zhida Sun, Zhenyao Zhang, Min Lu 0002, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CHI | 5 |
| 2026 | Palette Aligned Image DiffusionabstractAbstract We introduce the Palette‐Adapter , a novel method for conditioning text‐to‐image diffusion models on a user‐specified color palette. While palettes are a compact and intuitive tool widely used in creative workflows, they introduce significant ambiguity and instability when used for conditioning image generation. Our approach addresses this challenge by interpreting palettes as sparse histograms and introducing two scalar control parameters: histogram entropy and palette‐to‐histogram distance , which allow flexible control over the degree of palette adherence and color variation. We further introduce a negative histogram mechanism that allows users to suppress specific undesired hues, improving adherence to the intended palette under the standard classifier‐free guidance mechanism. To ensure broad generalization across the color space, we train on a carefully curated dataset with balanced coverage of rare and common colors. Our method enables stable, semantically coherent generation across a wide range of palettes and prompts. We evaluate our method qualitatively, quantitatively, and through a human evaluation, and show that it consistently outperforms existing approaches in achieving both strong palette adherence and high image quality. Elad Aharoni, Noy Porat, Dani Lischinski, Ariel Shamir |
Comput. Graph. Forum | 3 |
| 2026 | Story2Board: A Training-Free Approach for Expressive Visual StorytellingabstractAbstract We present Story2Board, a training‐free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing. To address this, we introduce a lightweight consistency framework composed of two components: Latent Panel Anchoring, which preserves a shared character reference across panels, and Reciprocal Attention Value Mixing, which softly blends visual features between token pairs with strong reciprocal attention. Together, these mechanisms enhance coherence without architectural changes or fine‐tuning, enabling state‐of‐the‐art diffusion models to generate visually diverse yet consistent storyboards. To structure generation, we use an off‐the‐shelf language model to convert free‐form stories into grounded panel‐level prompts. To evaluate, we propose the Rich Storyboard Benchmark , a suite of open‐domain narratives designed to assess layout diversity and background‐grounded storytelling, in addition to consistency. We also introduce a new Scene Diversity metric that quantifies spatial and pose variation across storyboards. Our qualitative and quantitative results, as well as a user study, show that Story2Board produces more dynamic, coherent, and narratively engaging storyboards than existing baselines. Project page : https://daviddinkevich.github.io/Story2Board/ David Dinkevich, Matan Levy, Omri Avrahami, Dvir Samuel, Dani Lischinski |
Comput. Graph. Forum | 5 |
| 2025 | Click2Mask: Local Editing with Dynamic Mask GenerationabstractRecent advancements in generative models have revolutionized image generation and editing, making these tasks accessible to non-experts. This paper focuses on local image editing, particularly the task of adding new content to a loosely specified area. Existing methods often require a precise mask or a detailed description of the location, which can be cumbersome and prone to errors. We propose Click2Mask, a novel approach that simplifies the local editing process by requiring only a single point of reference (in addition to the content description). A mask is dynamically grown around this point during a Blended Latent Diffusion (BLD) process, guided by a masked CLIP-based semantic loss. Click2Mask surpasses the limitations of segmentation-based and fine-tuning dependent methods, offering a more user-friendly and contextually accurate solution. Our experiments demonstrate that Click2Mask not only minimizes user effort but also enables competitive or superior local image manipulations compared to SoTA methods, according to both human judgement and automatic metrics. Key contributions include the simplification of user input, the ability to freely add objects unconstrained by existing segments, and the integration potential of our dynamic mask approach within other editing methods. Omer Regev, Omri Avrahami, Dani Lischinski |
AAAI | 3 |
| 2025 | EditInspector: A Benchmark for Evaluation of Text-Guided Image EditsabstractText-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their quality. To address this need, we introduce EditInspector, a novel benchmark for evaluation of text-guided image edits, based on human annotations collected using an extensive template for edit verification. We leverage EditInspector to evaluate the performance of state-of-the-art (SoTA) vision and language models in assessing edits across various dimensions, including accuracy, artifact detection, visual quality, seamless integration with the image scene, adherence to common sense, and the ability to describe edit-induced changes. Our findings indicate that current models struggle to evaluate edits comprehensively and frequently hallucinate when describing the changes. To address these challenges, we propose two novel methods that outperform SoTA models in both artifact detection and difference caption generation. Ron Yosef, Yonatan Bitton, Dani Lischinski, Moran Yanuka |
ACL (1) | 3 |
| 2025 | Creative Blends of Visual Concepts
Zhida Sun, Zhenyao Zhang, Min Lu 0002, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CHI | 5 |
| 2025 | EmoEdit: Evoking Emotions through Image ManipulationabstractAffective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotions. This task is inherently complex due to its twofold objective: evoking the intended emotion while preserving image composition. Existing AIM methods primarily adjust color and style, often failing to elicit precise, profound emotional shifts. Drawing on psychological insights, we introduce EmoEdit, which extends AIM by incorporating content modifications to enhance emotional impact. Specifically, we construct EmoEditSet, a large-scale AIM dataset of 40,120 paired data through emotion attribution and data construction. To make generative models emotion-aware, we design an Emotion Adapter and train it using EmoEditSet. We further propose an instruction loss to capture semantic variations in each data pair. Our method is evaluated both qualitatively and quantitatively, demonstrating superior performance over state-of-the-art techniques. Additionally, we showcase the portability of our Emotion Adapter to other diffusion-based models, enhancing their emotion knowledge with diverse semantics. Code is available at: https://github.com/JingyuanYY/EmoEdit. Jingyuan Yang 0002, Weibin Luo, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 4 |
| 2025 | Stable Flow: Vital Layers for Training-Free Image EditingabstractDiffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional UNet architecture with the Diffusion Transformer (DiT), and employed flow-matching for improved training and sampling. However, they exhibit limited generation diversity. In this work, we leverage this limitation to perform consistent image edits via selective injection of attention features. The main challenge is that, unlike the UNet-based models, DiT lacks a coarse-to-fine synthesis structure, making it unclear in which layers to perform the injection. Therefore, we propose an automatic method to identify "vital layers" within DiT, crucial for image formation, and demonstrate how these layers facilitate a range of controlled stable edits, from non-rigid modifications to object addition, using the same mechanism. Next, to enable real-image editing, we introduce an improved image inversion method for flow models. Finally, we evaluate our approach through qualitative and quantitative comparisons, along with a user study, and demonstrate its effectiveness across multiple applications. Omri Avrahami, Or Patashnik, Ohad Fried, Egor Nemchinov, Kfir Aberman, Dani Lischinski, Daniel Cohen-Or |
CVPR | 6 |
| 2025 | DiffTex: Differentiable Texturing for Architectural Proxy ModelsabstractSimplified proxy models are commonly used to represent architectural structures, reducing storage requirements and enabling real-time rendering. However, the geometric simplifications inherent in proxies result in a loss of fine color and geometric details, making it essential for textures to compensate for the loss. Preserving the rich texture information from the original dense architectural reconstructions remains a daunting task, particularly when working with unordered RGB photographs. We propose an automated method for generating realistic texture maps for architectural proxy models at the texel level from an unordered collection of registered photographs. Our approach establishes correspondences between texels on a UV map and pixels in the input images, with each texel's color computed as a weighted blend of associated pixel values. Using differentiable rendering, we optimize blending parameters to ensure photometric and perspective consistency, while maintaining seamless texture coherence. Experimental results demonstrate the effectiveness and robustness of our method across diverse architectural models and varying photographic conditions, enabling the creation of high-quality textures that preserve visual fidelity and structural detail. Weidan Xiong, Yongli Wu, Bochuan Zeng, Jianwei Guo 0003, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2024 | Data Roaming and Quality Assessment for Composed Image RetrievalabstractThe task of Composed Image Retrieval (CoIR) involves queries that combine image and text modalities, allowing users to express their intent more effectively. However, current CoIR datasets are orders of magnitude smaller compared to other vision and language (V&L) datasets. Additionally, some of these datasets have noticeable issues, such as queries containing redundant modalities. To address these shortcomings, we introduce the Large Scale Composed Image Retrieval (LaSCo) dataset, a new CoIR dataset which is ten times larger than existing ones. Pre-training on our LaSCo, shows a noteworthy improvement in performance, even in zero-shot. Furthermore, we propose a new approach for analyzing CoIR datasets and methods, which detects modality redundancy or necessity, in queries. We also introduce a new CoIR baseline, the Cross-Attention driven Shift Encoder (CASE). This baseline allows for early fusion of modalities using a cross-attention module and employs an additional auxiliary task during training. Our experiments demonstrate that this new baseline outperforms the current state-of-the-art methods on established benchmarks like FashionIQ and CIRR. Matan Levy, Rami Ben-Ari, Nir Darshan, Dani Lischinski |
AAAI | 4 |
| 2024 | Generating Non-Stationary Textures Using Self-RectificationabstractThis paper addresses the challenge of example-based non-stationary texture synthesis. We introduce a novel two-step approach wherein users first modify a reference texture using standard image editing tools, yielding an initial rough target for the synthesis. Subsequently, our proposed method, termed “self-rectification”, automatically refines this target into a coherent, seamless texture, while faithfully preserving the distinct visual characteristics of the reference exemplar. Our method leverages a pretrained diffusion network, and uses self-attention mechanisms, to grad-ually align the synthesized texture with the reference, en-suring the retention of the structures in the provided target. Through experimental validation, our approach ex-hibits exceptional proficiency in handling non-stationary textures, demonstrating significant advancements in texture synthesis when compared to existing state-of-the-art techniques. Code is available at https://github.com/xiaorongjun000/Self-Rectification Yang Zhou 0007, Rongjun Xiao, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 3 |
| 2024 | Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
Brian Gordon, Yonatan Bitton, Yonatan Shafir, Roopal Garg, Xi Chen 0071, Dani Lischinski, Daniel Cohen-Or, Idan Szpektor |
ECCV (57) | 6 |
| 2024 | Noise-free Score DistillationabstractScore Distillation Sampling (SDS) has emerged as the de facto approach for text-to-content generation in non-image domains. In this paper, we reexamine the SDS process and introduce a straightforward interpretation that demystifies the necessity for large Classifier-Free Guidance (CFG) scales, rooted in the distillation of an undesired noise term. Building upon our interpretation, we propose a novel Noise-Free Score Distillation (NFSD) process, which requires minimal modifications to the original SDS framework. Through this streamlined design, we achieve more effective distillation of pre-trained text-to-image diffusion models while using a nominal CFG scale. This strategic choice allows us to prevent the over-smoothing of results, ensuring that the generated data is both realistic and complies with the desired prompt. To demonstrate the efficacy of NFSD, we provide qualitative examples that compare NFSD and SDS, as well as several other methods. Oren Katzir, Or Patashnik, Daniel Cohen-Or, Dani Lischinski |
ICLR | 4 |
| 2024 | DiffUHaul: A Training-Free Method for Object Dragging in Images
Omri Avrahami, Rinon Gal, Gal Chechik, Ohad Fried, Dani Lischinski, Arash Vahdat, Weili Nie |
SIGGRAPH Asia | 5 |
| 2024 | CLIP-Flow: Decoding images encoded in CLIP spaceabstractThis study introduces CLIP-Flow, a novel network for generating images from a given image or text. To effectively utilize the rich semantics contained in both modalities, we designed a semantics-guided methodology for image- and text-to-image synthesis. In particular, we adopted Contrastive Language-Image Pretraining (CLIP) as an encoder to extract semantics and StyleGAN as a decoder to generate images from such information. Moreover, to bridge the embedding space of CLIP and latent space of StyleGAN, real NVP is employed and modified with activation normalization and invertible convolution. As the images and text in CLIP share the same representation space, text prompts can be fed directly into CLIP-Flow to achieve text-to-image synthesis. We conducted extensive experiments on several datasets to validate the effectiveness of the proposed image-to-image synthesis method. In addition, we tested on the public dataset Multi-Modal CelebA-HQ, for text-to-image synthesis. Experiments validated that our approach can generate high-quality text-matching images, and is comparable with state-of-the-art methods, both qualitatively and quantitatively. Jingyuan Yang 0002, Or Patashnik, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
Comput. Vis. Media | 5 |
| 2023 | SpaText: Spatio-Textual Representation for Controllable Image GenerationabstractRecent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous attempts to provide such controls were hindered by their reliance on a fixed set of labels. To this end, we present SpaText — a new method for text-to-image generation using open-vocabulary scene control. In addition to a global text prompt that describes the entire scene, the user provides a segmentation map where each region of interest is annotated by a free-form natural language description. Due to lack of large-scale datasets that have a detailed textual description for each region in the image, we choose to leverage the current large-scale text-to-image datasets and base our approach on a novel CLIP-based spatio-textual representation, and show its effectiveness on two state-of-the-art diffusion models: pixel-based and latent-based. In addition, we show how to extend the classifier-free guidance method in diffusion models to the multi-conditional case and present an alternative accelerated inference algorithm. Finally, we offer several automatic evaluation metrics and use them, in addition to FID scores and a user study, to evaluate our method and show that it achieves state-of-the-art results on image generation with free-form textual scene control. Omri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta, Yaniv Taigman, Devi Parikh, Dani Lischinski, Ohad Fried, Xi Yin 0001 |
CVPR | 7 |
| 2023 | EmoSet: A Large-scale Visual Emotion Dataset with Rich AttributesabstractVisual Emotion Analysis (VEA) aims at predicting people’s emotional responses to visual stimuli. This is a promising, yet challenging, task in affective computing, which has drawn increasing attention in recent years. Most of the existing work in this area focuses on feature design, while little attention has been paid to dataset construction. In this work, we introduce EmoSet, the first large-scale visual emotion dataset annotated with rich attributes, which is superior to existing datasets in four aspects: scale, annotation richness, diversity, and data balance. EmoSet comprises 3.3 million images in total, with 118,102 of these images carefully labeled by human annotators, making it five times larger than the largest existing dataset. EmoSet includes images from social networks, as well as artistic images, and it is well balanced between different emotion categories. Motivated by psychological studies, in addition to emotion category, each image is also annotated with a set of describable emotion attributes: brightness, colorfulness, scene type, object class, facial expression, and human action, which can help understand visual emotions in a precise and interpretable way. The relevance of these emotion attributes is validated by analyzing the correlations between them and visual emotion, as well as by designing an attribute module to help visual emotion recognition. We be lieve EmoSet will bring some key insights and encourage further research in visual emotion analysis and understanding. Project page: https://vcc.tech/EmoSet. Jingyuan Yang 0002, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ICCV | 4 |
| 2023 | Chatting Makes Perfect: Chat-based Image RetrievalabstractChats emerge as an effective user-friendly approach for information retrieval, and are successfully employed in many domains, such as customer service, healthcare, and finance. However, existing image retrieval approaches typically address the case of a single query-to-image round, and the use of chats for image retrieval has been mostly overlooked. In this work, we introduce ChatIR: a chat-based image retrieval system that engages in a conversation with the user to elicit information, in addition to an initial query, in order to clarify the user's search intent. Motivated by the capabilities of today's foundation models, we leverage Large Language Models to generate follow-up questions to an initial image description. These questions form a dialog with the user in order to retrieve the desired image from a large corpus. In this study, we explore the capabilities of such a system tested on a large dataset and reveal that engaging in a dialog yields significant gains in image retrieval. We start by building an evaluation pipeline from an existing manually generated dataset and explore different modules and training strategies for ChatIR. Our comparison includes strong baselines derived from related applications trained with Reinforcement Learning. Our system is capable of retrieving the target image from a pool of 50K images with over 78% success rate after 5 dialogue rounds, compared to 75% when questions are asked by humans, and 64% for a single shot text-to-image retrieval.
Extensive evaluations reveal the strong capabilities and examine the limitations of CharIR under different settings. Project repository is available at https://github.com/levymsn/ChatIR. Matan Levy, Rami Ben-Ari, Nir Darshan, Dani Lischinski |
NeurIPS | 4 |
| 2023 | Break-A-Scene: Extracting Multiple Concepts from a Single ImageabstractText-to-image model personalization aims to introduce a user-provided concept to the model, allowing its synthesis in diverse contexts. However, current methods primarily focus on the case of learning a single concept from multiple images with variations in backgrounds and poses, and struggle when adapted to a different scenario. In this work, we introduce the task of textual scene decomposition: given a single image of a scene that may contain several concepts, we aim to extract a distinct text token for each concept, enabling fine-grained control over the generated scenes. To this end, we propose augmenting the input image with masks that indicate the presence of target concepts. These masks can be provided by the user or generated automatically by a pre-trained segmentation model. We then present a novel two-phase customization process that optimizes a set of dedicated textual embeddings (handles), as well as the model weights, striking a delicate balance between accurately capturing the concepts and avoiding overfitting. We employ a masked diffusion loss to enable handles to generate their assigned concepts, complemented by a novel loss on cross-attention maps to prevent entanglement. We also introduce union-sampling, a training strategy aimed to improve the ability of combining multiple concepts in generated images. We use several automatic metrics to quantitatively compare our method against several baselines, and further affirm the results using a user study. Finally, we showcase several applications of our method. Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel Cohen-Or, Dani Lischinski |
SIGGRAPH Asia | 5 |
| 2023 | What's in a Decade? Transforming Faces Through TimeabstractAbstract How can one visually characterize photographs of people over time? In this work, we describe theFaces Through Timedataset, which contains over a thousand portrait images per decade from the 1880s to the present day. Using our new dataset, we devise a framework for resynthesizing portrait images across time, imagining how a portrait taken during a particular decade might have looked like had it been taken in other decades. Our framework optimizes a family of per‐decade generators that reveal subtle changes that differentiate decades—such as different hairstyles or makeup—while maintaining the identity of the input portrait. Experiments show that our method can more effectively resynthesizing portraits across time compared to state‐of‐the‐art image‐to‐image translation methods, as well as attribute‐based and language‐guided portrait editing models. Our code and data will be available at facesthroughtime.github.io. Eric Ming Chen, Jin Sun 0011, Apoorv Khandelwal 0001, Dani Lischinski, Noah Snavely, Hadar Averbuch-Elor |
Comput. Graph. Forum | 4 |
| 2023 | Blended Latent DiffusionabstractThe tremendous progress in neural image generation, coupled with the emergence of seemingly omnipotent vision-language models has finally enabled text-based interfaces for creating and editing images. Handling generic images requires a diverse underlying generative model, hence the latest works utilize diffusion models, which were shown to surpass GANs in terms of diversity. One major drawback of diffusion models, however, is their relatively slow inference time. In this paper, we present an accelerated solution to the task of local text-driven editing of generic images, where the desired edits are confined to a user-provided mask. Our solution leverages a text-to-image Latent Diffusion Model (LDM), which speeds up diffusion by operating in a lower-dimensional latent space and eliminating the need for resource-intensive CLIP gradient calculations at each diffusion step. We first enable LDM to perform local image edits by blending the latents at each step, similarly to Blended Diffusion. Next we propose an optimization-based solution for the inherent inability of LDM to accurately reconstruct images. Finally, we address the scenario of performing local edits using thin masks. We evaluate our method against the available baselines both qualitatively and quantitatively and demonstrate that in addition to being faster, it produces more precise results. Omri Avrahami, Ohad Fried, Dani Lischinski |
ACM Trans. Graph. | 3 |
| 2022 | Blended Diffusion for Text-driven Editing of Natural ImagesabstractNatural language offers a highly intuitive interface for image editing. In this paper, we introduce the first solution for performing local (region-based) edits in generic natural images, based on a natural language description along with an ROI mask. We achieve our goal by leveraging and combining a pretrained language-image model (CLIP), to steer the edit towards a user-provided text prompt, with a denoising diffusion probabilistic model (DDPM) to generate natural-looking results. To seamlessly fuse the edited region with the unchanged parts of the image, we spatially blend noised versions of the input image with the local text-guided diffusion latent at a progression of noise levels. In addition, we show that adding augmentations to the diffusion process mitigates adversarial results. We compare against several baselines and related methods, both qualitatively and quantitatively, and show that our method outperforms these solutions in terms of overall realism, ability to preserve the background and matching the text. Finally, we show several text-driven editing applications, including adding a new object to an image, removing/replacing/altering existing objects, background replacement, and image extrapolation. Omri Avrahami, Dani Lischinski, Ohad Fried |
CVPR | 2 |
| 2022 | ShapeFormer: Transformer-based Shape Completion via Sparse RepresentationabstractWe present ShapeFormer, a transformer-based network that produces a distribution of object completions, conditioned on incomplete, and possibly noisy, point clouds. The resultant distribution can then be sampled to generate likely completions, each exhibiting plausible shape details while being faithful to the input. To facilitate the use of transformers for 3D, we introduce a compact 3D representation, vector quantized deep implicit function (VQDIF), that utilizes spatial sparsity to represent a close approximation of a 3D shape by a short sequence of discrete variables. Experiments demonstrate that ShapeFormer outperforms prior art for shape completion from ambiguous partial inputs in terms of both completion quality and diversity. We also show that our approach effectively handles a variety of shape types, incomplete patterns, and real-world scans. Xingguang Yan, Liqiang Lin, Niloy J. Mitra, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 4 |
| 2022 | GAN Cocktail: Mixing GANs Without Dataset Access
Omri Avrahami, Dani Lischinski, Ohad Fried |
ECCV (23) | 2 |
| 2022 | Shape-Pose Disentanglement Using SE(3)-Equivariant Vector Neurons
Oren Katzir, Dani Lischinski, Daniel Cohen-Or |
ECCV (3) | 2 |
| 2022 | Classification-Regression for Chart Comprehension
Matan Levy, Rami Ben-Ari, Dani Lischinski |
ECCV (36) | 3 |
| 2022 | StyleAlign: Analysis and Applications of Aligned StyleGAN Models
Zongze Wu 0002, Yotam Nitzan, Eli Shechtman, Dani Lischinski |
ICLR | 4 |
| 2022 | DO-Conv: Depthwise Over-Parameterized Convolutional LayerabstractConvolutional layers are the core building blocks of Convolutional Neural Networks (CNNs). In this paper, we propose to augment a convolutional layer with an additional depthwise convolution, where each input channel is convolved with a different 2D kernel. The composition of the two convolutions constitutes an over-parameterization, since it adds learnable parameters, while the resulting linear operation can be expressed by a single convolution layer. We refer to this depthwise over-parameterized convolutional layer as DO-Conv, which is a novel way of over-parameterization. We show with extensive experiments that the mere replacement of conventional convolutional layers with DO-Conv layers boosts the performance of CNNs on many classical vision tasks, such as image classification, detection, and segmentation. Moreover, in the inference phase, the depthwise convolution is folded into the conventional convolution, reducing the computation to be exactly equivalent to that of a convolutional layer without over-parameterization. As DO-Conv introduces performance gains without incurring any computational complexity increase for inference, we advocate it as an alternative to the conventional convolutional layer. We open sourced an implementation of DO-Conv in Tensorflow, PyTorch and GluonCV at https://github.com/yangyanli/DO-Conv. Jinming Cao, Yangyan Li, Mingchao Sun, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen, Changhe Tu |
IEEE Trans. Image Process. | 5 |
| 2021 | StyleSpace Analysis: Disentangled Controls for StyleGAN Image GenerationabstractWe explore and analyze the latent style space of Style-GAN2, a state-of-the-art architecture for image generation, using models pretrained on several different datasets. We first show that StyleSpace, the space of channel-wise style parameters, is significantly more disentangled than the other intermediate latent spaces explored by previous works. Next, we describe a method for discovering a large collection of style channels, each of which is shown to control a distinct visual attribute in a highly localized and dis-entangled manner. Third, we propose a simple method for identifying style channels that control a specific attribute, using a pretrained classifier or a small number of example images. Manipulation of visual attributes via these StyleSpace controls is shown to be better disentangled than via those proposed in previous works. To show this, we make use of a newly proposed Attribute Dependency metric. Finally, we demonstrate the applicability of StyleSpace controls to the manipulation of real images. Our findings pave the way to semantically meaningful and well-disentangled image manipulations via simple and intuitive interfaces. Zongze Wu 0002, Dani Lischinski, Eli Shechtman |
CVPR | 2 |
| 2021 | ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic SegmentationabstractRGB-D semantic segmentation has attracted increasing attention over the past few years. Existing methods mostly employ homogeneous convolution operators to consume the RGB and depth features, ignoring their intrinsic differences. In fact, the RGB values capture the photometric appearance properties in the projected image space, while the depth feature encodes both the shape of a local geometry as well as the base (whereabout) of it in a larger context. Compared with the base, the shape probably is more inherent and has a stronger connection to the semantics, and thus is more critical for segmentation accuracy. Inspired by this observation, we introduce a Shape-aware Convolutional layer (ShapeConv) for processing the depth feature, where the depth feature is firstly decomposed into a shape-component and a base-component, next two learnable weights are introduced to cooperate with them independently, and finally a convolution is applied on the re-weighted combination of these two components. ShapeConv is model-agnostic and can be easily integrated into most CNNs to replace vanilla convolutional layers for semantic segmentation. Extensive experiments on three challenging indoor RGB-D semantic segmentation benchmarks, i.e., NYU-Dv2(-13,-40), SUN RGB-D, and SID, demonstrate the effectiveness of our ShapeConv when employing it over five popular architectures. Moreover, the performance of CNNs with ShapeConv is boosted without introducing any computation and memory increase in the inference phase. The reason is that the learnt weights for balancing the importance between the shape and base components in ShapeConv become constants in the inference phase, and thus can be fused into the following convolution, resulting in a network that is identical to one with vanilla convolutional layers. Jinming Cao, Hanchao Leng, Dani Lischinski, Daniel Cohen-Or, Changhe Tu, Yangyan Li |
ICCV | 3 |
| 2021 | StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryabstractInspired by the ability of StyleGAN to generate highly realistic images in a variety of domains, much recent work has focused on understanding how to use the latent spaces of StyleGAN to manipulate generated and real images. However, discovering semantically meaningful latent manipulations typically involves painstaking human examination of the many degrees of freedom, or an annotated collection of images for each desired manipulation. In this work, we explore leveraging the power of recently introduced Contrastive Language-Image Pre-training (CLIP) models in order to develop a text-based interface for StyleGAN image manipulation that does not require such manual effort. We first introduce an optimization scheme that utilizes a CLIP-based loss to modify an input latent vector in response to a user-provided text prompt. Next, we describe a latent mapper that infers a text-guided latent manipulation step for a given input image, allowing faster and more stable text-based manipulation. Finally, we present a method for mapping text prompts to input-agnostic directions in StyleGAN’s style space, enabling interactive text-driven image manipulation. Extensive results and comparisons demonstrate the effectiveness of our approaches. Or Patashnik, Zongze Wu 0002, Eli Shechtman, Daniel Cohen-Or, Dani Lischinski |
ICCV | 5 |
| 2021 | Fine-grained Foreground Retrieval via Teacher-Student LearningabstractForeground image retrieval is a challenging computer vision task. Given a background scene image with a bounding box indicating a target location, the goal is to retrieve a set of images of foreground objects from a given category, which are semantically compatible with the background. We formulate foreground retrieval as a self-supervised domain adaptation task, where the source domain consists of foreground images and the target domain of background images. Specifically, given pretrained object feature extraction networks that serve as teachers, we train a student network to infer compatible foreground features from background images. Thus, foregrounds and backgrounds are effectively mapped into a common feature space, enabling retrieval of the foregrounds that are closest to the target background in that space. A notable feature of our approach is that our training strategy does not require instance segmentation, unlike current state-of-the-art methods. Thus, our method may be applied to diverse foreground categories and background scene types and enables us to retrieve the foreground in a fine-grained manner, which is closer to the requirements of real world applications. Zongze Wu 0002, Dani Lischinski, Eli Shechtman |
WACV | 2 |
| 2021 | Weakly supervised 2D human pose transfer
Zhizhao Lin, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
Sci. China Inf. Sci. | 4 |
| 2021 | RGB×D: Learning depth-weighted RGB patches for RGB-D indoor semantic segmentation
Jinming Cao, Hanchao Leng, Daniel Cohen-Or, Dani Lischinski, Changhe Tu, Yangyan Li |
Neurocomputing | 4 |
| 2021 | MotioNet: 3D Human Motion Reconstruction from Monocular Video with Skeleton ConsistencyabstractWe introduce MotioNet , a deep neural network that directly reconstructs the motion of a 3D human skeleton from a monocular video. While previous methods rely on either rigging or inverse kinematics (IK) to associate a consistent skeleton with temporally coherent joint rotations, our method is the first data-driven approach that directly outputs a kinematic skeleton, which is a complete, commonly used motion representation. At the crux of our approach lies a deep neural network with embedded kinematic priors, which decomposes sequences of 2D joint positions into two separate attributes: a single, symmetric skeleton encoded by bone lengths, and a sequence of 3D joint rotations associated with global root positions and foot contact labels. These attributes are fed into an integrated forward kinematics (FK) layer that outputs 3D positions, which are compared to a ground truth. In addition, an adversarial loss is applied to the velocities of the recovered rotations to ensure that they lie on the manifold of natural joint rotations. The key advantage of our approach is that it learns to infer natural joint rotations directly from the training data rather than assuming an underlying model, or inferring them from joint positions using a data-agnostic IK solver. We show that enforcing a single consistent skeleton along with temporally coherent joint rotations constrains the solution space, leading to a more robust handling of self-occlusions and depth ambiguities. Mingyi Shi, Kfir Aberman, Andreas Aristidou, Taku Komura, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2021 | Palettailor: Discriminable Colorization for Categorical DataabstractWe present an integrated approach for creating and assigning color palettes to different visualizations such as multi-class scatterplots, line, and bar charts. While other methods separate the creation of colors from their assignment, our approach takes data characteristics into account to produce color palettes, which are then assigned in a way that fosters better visual discrimination of classes. To do so, we use a customized optimization based on simulated annealing to maximize the combination of three carefully designed color scoring functions: point distinctness, name difference, and color discrimination. We compare our approach to state-of-the-art palettes with a controlled user study for scatterplots and line charts, furthermore we performed a case study. Our results show that Palettailor, as a fully-automated approach, generates color palettes with a higher discrimination quality than existing approaches. The efficiency of our optimization allows us also to incorporate user modifications into the color selection process. Kecheng Lu 0002, Mi Feng, Xin Chen 0075, Michael Sedlmair, Oliver Deussen, Dani Lischinski, Zhanglin Cheng, Yunhai Wang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Cross-Domain Cascaded Deep Translation
Oren Katzir, Dani Lischinski, Daniel Cohen-Or |
ECCV (2) | 2 |
| 2020 | CrossNet: Latent Cross-Consistency for Unpaired Image TranslationabstractRecent GAN-based architectures have been able to deliver impressive performance on the general task of image-to-image translation. In particular, it was shown that a wide variety of image translation operators may be learned from two image sets, containing images from two different domains, without establishing an explicit pairing between the images. This was made possible by introducing clever regularizers to overcome the under-constrained nature of the unpaired translation problem. In this work, we introduce a novel architecture for unpaired image translation, and explore several new regularizes enabled by it. Specifically, our architecture comprises a pair of GANs, as well as a pair of translators between their respective latent spaces. These cross-translators enable us to impose several regularizing constraints on the learnt image translation operator, collectively referred to as latent cross-consistency. Our results show that our proposed architecture and latent cross-consistency constraints are able to outperform the existing state-of-the-art on a variety of image translation tasks. Omry Sendik, Dani Lischinski, Daniel Cohen-Or |
WACV | 2 |
| 2020 | Fabricable dihedral Escher tessellations
Lin Lu 0001, Andrei Sharf, Dani Lischinski, Changhe Tu |
Comput. Aided Des. | 5 |
| 2020 | Fabricable Unobtrusive 3D-QR-Codes with Directional LightabstractAbstract QR code is a 2D matrix barcode widely used for product tracking, identification, document management and general marketing. Recently, there have been various attempts to utilize QR codes in 3D manufacturing by carving QR codes on the surface of the printed 3D shape. Nevertheless, significant shape editing and modulation may be required to allow readability of the embedded 3D‐QR‐codes with good decoding accuracy. In this paper, we introduce a novel QR code 3D fabrication framework aimed at unobtrusive embedding of 3D‐QR‐codes in the shape hence introducing minimal shape modulation. Essentially, our method computes bi‐directional carvings in the 3D shape surface to obtain the black‐and‐white QR pattern. By using a directional light source, the black‐and‐white QR pattern emerges as lighted and shadow casted blocks on the shape respectively. To account for minimal modulation and elusiveness, we optimize the QR code carving w.r.t. shape geometry, visual disparity and light source position. Our technique employs a simulation of lighting phenomena through carved modules on the shape to ensure adequate contrast of the printed 3D‐QR‐code. Hao Peng 0001, Peiqing Liu, Lin Lu 0001, Andrei Sharf, Dani Lischinski, Baoquan Chen |
Comput. Graph. Forum | 6 |
| 2020 | Skeleton-aware networks for deep motion retargetingabstractWe introduce a novel deep learning framework for data-driven motion retargeting between skeletons, which may have different structure, yet corresponding to homeomorphic graphs. Importantly, our approach learns how to retarget without requiring any explicit pairing between the motions in the training set. We leverage the fact that different homeomorphic skeletons may be reduced to a common primal skeleton by a sequence of edge merging operations, which we refer to as skeletal pooling. Thus, our main technical contribution is the introduction of novel differentiable convolution, pooling, and unpooling operators. These operators are skeleton-aware , meaning that they explicitly account for the skeleton's hierarchical structure and joint adjacency, and together they serve to transform the original motion into a collection of deep temporal features associated with the joints of the primal skeleton. In other words, our operators form the building blocks of a new deep motion processing framework that embeds the motion into a common latent space, shared by a collection of homeomorphic skeletons. Thus, retargeting can be achieved simply by encoding to, and decoding from this latent space. Our experiments show the effectiveness of our framework for motion retargeting, as well as motion processing in general, compared to existing approaches. Our approach is also quantitatively evaluated on a synthetic dataset that contains pairs of motions applied to different skeletons. To the best of our knowledge, our method is the first to perform retargeting between skeletons with differently sampled kinematic chains, without any paired examples. Kfir Aberman, Peizhuo Li, Dani Lischinski, Olga Sorkine-Hornung, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 3 |
| 2020 | Unpaired motion style transfer from video to animationabstractTransferring the motion style from one animation clip to another, while preserving the motion content of the latter, has been a long-standing problem in character animation. Most existing data-driven approaches are supervised and rely on paired data, where motions with the same content are performed in different styles. In addition, these approaches are limited to transfer of styles that were seen during training. In this paper, we present a novel data-driven framework for motion style transfer, which learns from an unpaired collection of motions with style labels, and enables transferring motion styles not observed during training. Furthermore, our framework is able to extract motion styles directly from videos, bypassing 3D reconstruction, and apply them to the 3D input motion. Our style transfer network encodes motions into two latent codes, for content and for style, each of which plays a different role in the decoding (synthesis) process. While the content code is decoded into the output motion by several temporal convolutional layers, the style code modifies deep features via temporally invariant adaptive instance normalization (AdaIN). Moreover, while the content code is encoded from 3D joint rotations, we learn a common embedding for style from either 3D or 2D joint positions, enabling style extraction from videos. Our results are comparable to the state-of-the-art, despite not requiring paired training data, and outperform other methods when transferring previously unseen styles. To our knowledge, we are the first to demonstrate style transfer directly from videos to 3D animations - an ability which enables one to extend the set of style examples far beyond motions captured by MoCap systems. Kfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 3 |
| 2020 | Inverse Procedural Modeling of Branching Structures by Inferring L-SystemsabstractWe introduce an inverse procedural modeling approach that learns L-system representations of pixel images with branching structures. Our fully automatic model generates a compact set of textual rewriting rules that describe the input. We use deep learning to discover atomic structures such as line segments or branchings. Orientation and scaling of these structures are determined and the detected structures are combined into a tree. The initial representation is analyzed, and repeating parts are encoded into a small grammar by using greedy optimization while the user can control the size of the detected rules. The output is an L-system that represents the input image as a simple text and a set of terminal symbols. We apply our approach to a variety of examples, demonstrate its robustness against noise and blur, and we show that it can detect user sketches and complex input structures. Jianwei Guo 0003, Haiyong Jiang, Bedrich Benes, Oliver Deussen, Xiaopeng Zhang 0001, Dani Lischinski, Hui Huang 0004 |
ACM Trans. Graph. | 6 |
| 2020 | Differentiable refraction-tracing for mesh reconstruction of transparent objectsabstractCapturing the 3D geometry of transparent objects is a challenging task, ill-suited for general-purpose scanning and reconstruction techniques, since these cannot handle specular light transport phenomena. Existing state-of-the-art methods, designed specifically for this task, either involve a complex setup to reconstruct complete refractive ray paths, or leverage a data-driven approach based on synthetic training data. In either case, the reconstructed 3D models suffer from over-smoothing and loss of fine detail. This paper introduces a novel, high precision, 3D acquisition and reconstruction method for solid transparent objects. Using a static background with a coded pattern, we establish a mapping between the camera view rays and locations on the background. Differentiable tracing of refractive ray paths is then used to directly optimize a 3D mesh approximation of the object, while simultaneously ensuring silhouette consistency and smoothness. Extensive experiments and comparisons demonstrate the superior accuracy of our method. Jiahui Lyu, Bojian Wu, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 3 |
| 2020 | Unsupervised K-modal styled content generationabstractThe emergence of deep generative models has recently enabled the automatic generation of massive amounts of graphical content, both in 2D and in 3D. Generative Adversarial Networks (GANs) and style control mechanisms, such as Adaptive Instance Normalization (AdaIN), have proved particularly effective in this context, culminating in the state-of-the-art StyleGAN architecture. While such models are able to learn diverse distributions, provided a sufficiently large training set, they are not well-suited for scenarios where the distribution of the training data exhibits a multi-modal behavior. In such cases, reshaping a uniform or normal distribution over the latent space into a complex multi-modal distribution in the data domain is challenging, and the generator might fail to sample the target distribution well. Furthermore, existing unsupervised generative models are not able to control the mode of the generated samples independently of the other visual attributes, despite the fact that they are typically disentangled in the training data. In this paper, we introduce uMM-GAN, a novel architecture designed to better model multi-modal distributions, in an unsupervised fashion. Building upon the StyleGAN architecture, our network learns multiple modes, in a completely unsupervised manner , and combines them using a set of learned weights. We demonstrate that this approach is capable of effectively approximating a complex distribution as a superposition of multiple simple ones. We further show that uMM-GAN effectively disentangles between modes and style, thereby providing an independent degree of control over the generated content. Omry Sendik, Dani Lischinski, Daniel Cohen-Or |
ACM Trans. Graph. | 2 |
| 2019 | ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object SegmentationabstractMulti-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a one-directional, top-down and bottom-up, manner. In this work, we introduce ZigZagNet, which aggregates a richer multi-context feature map by using not only dense top-down and bottom-up propagation, but also by introducing pathways crossing between different levels of the top-down and the bottom-up hierarchies, in a zig-zag fashion. Furthermore, the context information is exchanged and aggregated over multiple stages, where the fused feature maps from one stage are fed into the next one, yielding a more comprehensive context for improved segmentation performance. Our extensive evaluation on the public benchmarks demonstrates that ZigZagNet surpasses the state-of-the-art accuracy for both semantic segmentation and instance segmentation tasks. Di Lin 0002, Dingguo Shen, Siting Shen, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 5 |
| 2019 | Deep Video-Based Performance CloningabstractAbstract We present a new video‐based performance cloning technique. After training a deep generative network using a reference video capturing the appearance and dynamics of a target actor, we are able to generate videos where this actor reenacts other performances. All of the training data and the driving performances are provided as ordinary video segments, without motion capture or depth information. Our generative model is realized as a deep neural network with two branches, both of which train the same space‐time conditional generator, using shared weights. One branch, responsible for learning to generate the appearance of the target actor in various poses, uses paired training data, self‐generated from the reference video. The second branch uses unpaired data to improve generation of temporally coherent video renditions of unseen pose sequences. Through data augmentation, our network is able to synthesize images of the target actor in poses never captured by the reference video. We demonstrate a variety of promising results, where our method is able to generate temporally coherent videos, for challenging scenarios where the reference and driving videos consist of very different dance performances. Kfir Aberman, Mingyi Shi, Jing Liao 0001, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
Comput. Graph. Forum | 4 |
| 2019 | What's in a Face? Metric Learning for Face CharacterizationabstractAbstract We present a method for determining which facial parts (mouth, nose, etc.) best characterize an individual, given a set of that individual's portraits. We introduce a novel distinctiveness analysis of a set of portraits, which leverages the deep features extracted by a pre‐trained face recognition CNN and a hair segmentation FCN, in the context of a weakly supervised metric learning scheme. Our analysis enables the generation of a polarized class activation map (PCAM) for an individual's portrait via a transformation that localizes and amplifies the discriminative regions of the deep feature maps extracted by the aforementioned networks. A user study that we conducted shows that there is a surprisingly good agreement between the face parts that users indicate as characteristic and the face parts automatically selected by our method. We demonstrate a few applications of our method, including determining the most and the least representative portraits among a set of portraits of an individual, and the creation of facial hybrids: portraits that combine the characteristic recognizable facial features of two individuals. Our face characterization analysis is also effective for ranking portraits in order to find an individual's look‐alikes (Doppelgängers). Omry Sendik, Dani Lischinski, Daniel Cohen-Or |
Comput. Graph. Forum | 2 |
| 2019 | Point Pattern Synthesis via Irregular ConvolutionabstractAbstract Point pattern synthesis is a fundamental tool with various applications in computer graphics. To synthesize a point pattern, some techniques have taken an example‐based approach, where the user provides a small exemplar of the target pattern. However, it remains challenging to synthesize patterns that faithfully capture the structures in the given exemplar. In this paper, we present a new example‐based point pattern synthesis method that preserves both local and non‐local structures present in the exemplar. Our method leverages recent neural texture synthesis techniques that have proven effective in synthesizing structured textures. The network that we present is end‐to‐end. It utilizes an irregular convolution layer, which converts a point pattern into a gridded feature map, to directly optimize point coordinates. The synthesis is then performed by matching inter‐ and intra‐correlations of the responses produced by subsequent convolution layers. We demonstrate that our point pattern synthesis qualitatively outperforms state‐of‐the‐art methods on challenging structured patterns, and enables various graphical applications, such as object placement in natural scenes, creative element patterns or realistic urban layouts in a 3D virtual environment. Peihan Tu, Dani Lischinski, Hui Huang 0004 |
Comput. Graph. Forum | 2 |
| 2019 | Learning character-agnostic motion for motion retargeting in 2DabstractAnalyzing human motion is a challenging task with a wide variety of applications in computer vision and in graphics. One such application, of particular importance in computer animation, is the retargeting of motion from one performer to another. While humans move in three dimensions, the vast majority of human motions are captured using video, requiring 2D-to-3D pose and camera recovery, before existing retargeting approaches may be applied. In this paper, we present a new method for retargeting video-captured motion between different human performers, without the need to explicitly reconstruct 3D poses and/or camera parameters. In order to achieve our goal, we learn to extract, directly from a video, a high-level latent motion representation, which is invariant to the skeleton geometry and the camera view. Our key idea is to train a deep neural network to decompose temporal sequences of 2D poses into three components: motion, skeleton, and camera view-angle. Having extracted such a representation, we are able to re-combine motion with novel skeletons and camera views, and decode a retargeted temporal sequence, which we compare to a ground truth from a synthetic dataset. We demonstrate that our framework can be used to robustly extract human motion from videos, bypassing 3D reconstruction, and outperforming existing retargeting methods, when applied to videos in-the-wild. It also enables additional applications, such as performance cloning, video-driven cartoons, and motion retrieval. Kfir Aberman, Rundi Wu, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 3 |
| 2019 | SAGNet: structure-aware generative network for 3D-shape modelingabstractWe present SAGNet, a structure-aware generative model for 3D shapes. Given a set of segmented objects of a certain class, the geometry of their parts and the pairwise relationships between them (the structure) are jointly learned and embedded in a latent space by an autoencoder. The encoder intertwines the geometry and structure features into a single latent code, while the decoder disentangles the features and reconstructs the geometry and structure of the 3D model. Our autoencoder consists of two branches, one for the structure and one for the geometry. The key idea is that during the analysis, the two branches exchange information between them, thereby learning the dependencies between structure and geometry and encoding two augmented features, which are then fused into a single latent code. This explicit intertwining of information enables separately controlling the geometry and the structure of the generated models. We evaluate the performance of our method and conduct an ablation study. We explicitly show that encoding of shapes accounts for both similarities in structure and geometry. A variety of quality results generated by SAGNet are presented. Di Lin 0002, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2018 | Multi-scale Context Intertwining for Semantic Segmentation
Di Lin 0002, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ECCV (3) | 3 |
| 2018 | Fast Penetration Volume for Rigid BodiesabstractAbstract Handling collisions among a large number of bodies can be a performance bottleneck in video games and many other real‐time applications. We present a new framework for detecting and resolving collisions using the penetration volume as an interpenetration measure. Given two non‐convex polyhedral bodies, a new sampling paradigm locates their near‐contact configurations in advance, and stores associated contact information in a compact database. At runtime, we retrieve a given configuration's nearest neighbors. By taking advantage of the penetration volume's continuity, cheap geometric methods can use the neighbors to estimate contact information as well as a translational gradient. This results in an extremely fast, geometry‐independent, and trivially parallelizable computation, which constitutes the first global volume‐based collision resolution. When processing multiple collisions simultaneously on a 4‐core processor, the average running cost is as low as5 μs. Furthermore, no additional proximity or contact‐regions queries are required. These results are orders of magnitude faster than previous penetration volume approaches. D. Nirel, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2018 | Neural best-buddies: sparse cross-domain correspondenceabstractCorrespondence between images is a fundamental problem in computer vision, with a variety of graphics applications. This paper presents a novel method for sparse cross-domain correspondence. Our method is designed for pairs of images where the main objects of interest may belong to different semantic categories and differ drastically in shape and appearance, yet still contain semantically related or geometrically similar parts. Our approach operates on hierarchies of deep features, extracted from the input images by a pre-trained CNN. Specifically, starting from the coarsest layer in both hierarchies, we search for Neural Best Buddies (NBB): pairs of neurons that are mutual nearest neighbors. The key idea is then to percolate NBBs through the hierarchy, while narrowing down the search regions at each level and retaining only NBBs with significant activations. Furthermore, in order to overcome differences in appearance, each pair of search regions is transformed into a common appearance. We evaluate our method via a user study, in addition to comparisons with alternative correspondence approaches. The usefulness of our method is demonstrated using a variety of graphics applications, including cross-domain image alignment, creation of hybrid images, automatic image morphing, and more. Kfir Aberman, Jing Liao 0001, Mingyi Shi, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2018 | Appearance Modeling via Proxy-to-Image AlignmentabstractEndowing 3D objects with realistic surface appearance is a challenging and time-demanding task, as real-world surfaces typically exhibit a plethora of spatially variant geometric and photometric detail. Not surprisingly, computer artists commonly use images of real-world objects as an inspiration and a reference for their digital creations. However, despite two decades of research on image-based modeling, there are still no tools available for automatically extracting the detailed appearance (microgeometry and texture) of a 3D surface from a single image. In this article, we present a novel user-assisted approach for quickly and easily extracting a nonparametric appearance model from a single photograph of a reference object. The extraction process requires a user-provided proxy, whose geometry roughly approximates that of the object in the image. Since the proxy is just a rough approximation, it is necessary to align and deform it so as to match the reference object. The main contribution of this work is a novel technique to perform such an alignment, which enables accurate joint recovery of geometric detail and reflectance. The correlations between the recovered geometry at various scales and the spatially varying reflectance constitute a nonparametric appearance model. Once extracted, the appearance model may then be applied to various 3D shapes, whose large-scale geometry may differ considerably from that of the original reference object. Thus, our approach makes it possible to construct an appearance library, allowing users to easily enrich detail-less 3D shapes with realistic geometric detail and surface texture. Hui Huang 0004, Ke Xie 0001, Dani Lischinski, Minglun Gong, Xin Tong 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2018 | Creating and chaining camera moves for quadrotor videographyabstractCapturing aerial videos with a quadrotor-mounted camera is a challenging creative task, as it requires the simultaneous control of the quadrotor's motion and the mounted camera's orientation. Letting the drone follow a pre-planned trajectory is a much more appealing option, and recent research has proposed a number of tools designed to automate the generation of feasible camera motion plans; however, these tools typically require the user to specify and edit the camera path, for example by providing a complete and ordered sequence of key viewpoints. In this paper, we propose a higher level tool designed to enable even novice users to easily capture compelling aerial videos of large-scale outdoor scenes. Using a coarse 2.5D model of a scene, the user is only expected to specify starting and ending viewpoints and designate a set of landmarks, with or without a particular order. Our system automatically generates a diverse set of candidate local camera moves for observing each landmark, which are collision-free, smooth, and adapted to the shape of the landmark. These moves are guided by a landmark-centric view quality field, which combines visual interest and frame composition. An optimal global camera trajectory is then constructed that chains together a sequence of local camera moves, by choosing one move for each landmark and connecting them with suitable transition trajectories. This task is formulated and solved as an instance of the Set Traveling Salesman Problem. Ke Xie 0001, Shengqiu Huang, Dani Lischinski, Marc Christie, Kai Xu 0004, Minglun Gong, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2018 | Non-stationary texture synthesis by adversarial expansionabstractThe real world exhibits an abundance of non-stationary textures. Examples include textures with large scale structures, as well as spatially variant and inhomogeneous textures. While existing example-based texture synthesis methods can cope well with stationary textures, non-stationary textures still pose a considerable challenge, which remains unresolved. In this paper, we propose a new approach for example-based non-stationary texture synthesis. Our approach uses a generative adversarial network (GAN), trained to double the spatial extent of texture blocks extracted from a specific texture exemplar. Once trained, the fully convolutional generator is able to expand the size of the entire exemplar, as well as of any of its sub-blocks. We demonstrate that this conceptually simple approach is highly effective for capturing large scale structures, as well as other non-stationary attributes of the input exemplar. As a result, it can cope with challenging textures, which, to our knowledge, no other existing method can handle. Yang Zhou 0007, Zhen Zhu 0006, Xiang Bai, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2017 | Joint Bi-layer Optimization for Single-Image Rain Streak RemovalabstractWe present a novel method for removing rain streaks from a single input image by decomposing it into a rain-free background layer B and a rain-streak layer R. A joint optimization process is used that alternates between removing rain-streak details from B and removing non-streak details from R. The process is assisted by three novel image priors. Observing that rain streaks typically span a narrow range of directions, we first analyze the local gradient statistics in the rain image to identify image regions that are dominated by rain streaks. From these regions, we estimate the dominant rain streak direction and extract a collection of rain-dominated patches. Next, we define two priors on the background layer B, one based on a centralized sparse representation and another based on the estimated rain direction. A third prior is defined on the rain-streak layer R, based on similarity of patches to the extracted rain patches. Both visual and quantitative comparisons demonstrate that our method outperforms the state-of-the-art. Lei Zhu 0003, Chi-Wing Fu, Dani Lischinski, Pheng-Ann Heng |
ICCV | 3 |
| 2017 | Analysis and Controlled Synthesis of Inhomogeneous TexturesabstractMany interesting real-world textures are inhomogeneous and/or anisotropic. An inhomogeneous texture is one where various visual properties exhibit significant changes across the texture's spatial domain. Examples include perceptible changes in surface color, lighting, local texture pattern and/or its apparent scale, and weathering effects, which may vary abruptly, or in a continuous fashion. An anisotropic texture is one where the local patterns exhibit a preferred orientation, which also may vary across the spatial domain. While many example-based texture synthesis methods can be highly effective when synthesizing uniform (stationary) isotropic textures, synthesizing highly non-uniform textures, or ones with spatially varying orientation, is a considerably more challenging task, which so far has remained underexplored. In this paper, we propose a new method for automatic analysis and controlled synthesis of such textures. Given an input texture exemplar, our method generates a source guidance map comprising: (i) a scalar progression channel that attempts to capture the low frequency spatial changes in color, lighting, and local pattern combined, and (ii) a direction field that captures the local dominant orientation of the texture. Having augmented the texture exemplar with this guidance map, users can exercise better control over the synthesized result by providing easily specified target guidance maps, which are used to constrain the synthesis process. Yang Zhou 0007, Huajie Shi, Dani Lischinski, Minglun Gong, Johannes Kopf 0001, Hui Huang 0004 |
Comput. Graph. Forum | 3 |
| 2016 | Synthesizing Training Images for Boosting Human 3D Pose EstimationabstractHuman 3D pose estimation from a single image is a challenging task with numerous applications. Convolutional Neural Networks (CNNs) have recently achieved superior performance on the task of 2D pose estimation from a single image, by training on images with 2D annotations collected by crowd sourcing. This suggests that similar success could be achieved for direct estimation of 3D poses. However, 3D poses are much harder to annotate, and the lack of suitable annotated training images hinders attempts towards end-to-end solutions. To address this issue, we opt to automatically synthesize training images with ground truth pose annotations. Our work is a systematic study along this road. We find that pose space coverage and texture diversity are the key ingredients for the effectiveness of synthetic training data. We present a fully automatic, scalable approach that samples the human pose space for guiding the synthesis procedure and extracts clothing textures from real images. Furthermore, we explore domain adaptation for bridging the gap between our synthetic training images and real testing photos. We demonstrate that CNNs trained with our synthetic images out-perform those trained with real photos on 3D pose estimation tasks. Wenzheng Chen, Yangyan Li, Hao Su 0001, Zhenhua Wang 0002, Changhe Tu, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
3DV | 7 |
| 2016 | A Holistic Approach for Data-Driven Object Cutout
Huayong Xu, Yangyan Li, Wenzheng Chen, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACCV (1) | 4 |
| 2016 | Trip Synopsis: 60km in 60secabstractAbstract Computerized route planning tools are widely used today by travelers all around the globe, while 3D terrain and urban models are becoming increasingly elaborate and abundant. This makes it feasible to generate a virtual 3D flyby along a planned route. Such a flyby may be useful, either as a preview of the trip, or as an after‐the‐fact visual summary. However, a naively generated preview is likely to contain many boring portions, while skipping too quickly over areas worthy of attention. In this paper, we introduce 3D trip synopsis: a continuous visual summary of a trip that attempts to maximize the total amount of visual interest seen by the camera. The main challenge is to generate a synopsis of a prescribed short duration, while ensuring a visually smooth camera motion. Using an application‐specific visual interest metric, we measure the visual interest at a set of viewpoints along an initial camera path, and maximize the amount of visual interest seen in the synopsis by varying the speed along the route. A new camera path is then computed using optimization to simultaneously satisfy requirements, such as smoothness, focus and distance to the route. The process is repeated until convergence. The main technical contribution of this work is a new camera control method, which iteratively adjusts the camera trajectory and determines all of the camera trajectory parameters, including the camera position, altitude, heading, and tilt. Our results demonstrate the effectiveness of our trip synopses, compared to a number of alternatives. Hui Huang 0004, Dani Lischinski, Hao (Richard) Zhang, Minglun Gong, Marc Christie, Daniel Cohen-Or |
Comput. Graph. Forum | 2 |
| 2016 | Printed Perforated Lampshades for Continuous Projective ImagesabstractWe present a technique for designing three-dimensional- (3D) printed perforated lampshades that project continuous grayscale images onto the surrounding walls. Given the geometry of the lampshade and a target grayscale image, our method computes a distribution of tiny holes over the shell, such that the combined footprints of the light emanating through the holes form the target image on a nearby diffuse surface. Our objective is to approximate the continuous tones and the spatial detail of the target image to the extent possible within the constraints of the fabrication process. To ensure structural integrity, there are lower bounds on the thickness of the shell, the radii of the holes, and the minimal distances between adjacent holes. Thus, the holes are realized as thin tubes distributed over the lampshade surface. The amount of light passing through a single tube may be controlled by the tube’s radius and by its orientation (tilt angle). The core of our technique thus consists of determining a suitable configuration of the tubes: their distribution across the relevant portion of the lampshade, as well as the parameters (radius, tilt angle) of each tube. This is achieved by computing a capacity-constrained Voronoi tessellation over a suitably defined density function and embedding a tube inside the maximal inscribed circle of each tessellation cell. Haisen Zhao, Lin Lu 0001, Dani Lischinski, Andrei Sharf, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2015 | Self Tuning Texture OptimizationabstractAbstract The goal of example‐based texture synthesis methods is to generate arbitrarily large textures from limited exemplars in order to fit the exact dimensions and resolution required for a specific modeling task. The challenge is to faithfully capture all of the visual characteristics of the exemplar texture, without introducing obvious repetitions or unnatural looking visual elements. While existing non‐parametric synthesis methods have made remarkable progress towards this goal, most such methods have been demonstrated only on relatively low‐resolution exemplars. Real‐world high resolution textures often contain texture details at multiple scales, which these methods have difficulty reproducing faithfully. In this work, we present a new general‐purpose and fully automatic self‐tuning non‐parametric texture synthesis method that extends Texture Optimization by introducing several key improvements that result in superior synthesis ability. Our method is able to self‐tune its various parameters and weights and focuses on addressing three challenging aspects of texture synthesis: (i) irregular large scale structures are faithfully reproduced through the use of automatically generated and weighted guidance channels; (ii) repetition and smoothing of texture patches is avoided by new spatial uniformity constraints; (iii) a smart initialization strategy is used in order to improve the synthesis of regular and near‐regular textures, without affecting textures that do not exhibit regularities. We demonstrate the versatility and robustness of our completely automatic approach on a variety of challenging high‐resolution texture exemplars. Alexandre Kaspar, Boris Neubert, Dani Lischinski, Mark Pauly, Johannes Kopf 0001 |
Comput. Graph. Forum | 3 |
| 2015 | Hallucinating Stereoscopy from a Single ImageabstractAbstract We introduce a novel method for enabling stereoscopic viewing of a scene from a single pre‐segmented image. Rather than attempting full 3D reconstruction or accurate depth map recovery, we hallucinate a rough approximation of the scene's 3D model using a number of simple depth and occlusion cues and shape priors. We begin by depth‐sorting the segments, each of which is assumed to represent a separate object in the scene, resulting in a collection of depth layers. The shapes and textures of the partially occluded segments are then completed using symmetry and convexity priors. Next, each completed segment is converted to a union of generalized cylinders yielding a rough 3D model for each object. Finally, the object depths are refined using an iterative ground fitting process. The hallucinated 3D model of the scene may then be used to generate a stereoscopic image pair, or to produce images from novel viewpoints within a small neighborhood of the original view. Despite the simplicity of our approach, we show that it compares favorably with state‐of‐the‐art depth ordering methods. A user study was conducted showing that our method produces more convincing stereoscopic images than existing semi‐interactive and automatic single image depth recovery methods. Qiong Zeng, Wenzheng Chen, Changhe Tu, Daniel Cohen-Or, Dani Lischinski, Baoquan Chen |
Comput. Graph. Forum | 6 |
| 2015 | JumpCut: non-successive mask transfer and interpolation for video cutoutabstractWe introduce JumpCut, a new mask transfer and interpolation method for interactive video cutout. Given a source frame for which a foreground mask is already available, we compute an estimate of the foreground mask at another, typically non-successive, target frame. Observing that the background and foreground regions typically exhibit different motions, we leverage these differences by computing two separate nearest-neighbor fields (split-NNF) from the target to the source frame. These NNFs are then used to jointly predict a coherent labeling of the pixels in the target frame. The same split-NNF is also used to aid a novel edge classifier in detecting silhouette edges (S-edges) that separate the foreground from the background. A modified level set method is then applied to produce a clean mask, based on the pixel labels and the S-edges computed by the previous two steps. The resulting mask transfer method may also be used for coherently interpolating the foreground masks between two distant source frames. Our results demonstrate that the proposed method is significantly more accurate than the existing state-of-the-art on a wide variety of video sequences. Thus, it reduces the required amount of user effort, and provides a basis for an effective interactive video object cutout tool. Qingnan Fan, Fan Zhong 0001, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 3 |
| 2014 | Collaborative Personalization of Image Enhancement
Ashish Kapoor, Juan C. Caicedo, Dani Lischinski, Sing Bing Kang |
Int. J. Comput. Vis. | 3 |
| 2014 | Slippage-free background replacement for hand-held videoabstractWe introduce a method for replacing the background in a video of a moving foreground subject, when both the source video capturing the subject, and the target video capturing the new background scene, are natural videos, casually captured using a freely moving hand-held camera. We assume that the foreground subject has already been extracted, and focus on the challenging task of generating a video with a new background, such that the new background motion appears compatible with the original one. Failure to match the motion results in disturbing slippage or moonwalk artifacts, where the subject's feet appear to slide or slip over the ground. While matching the motion across the entire frame is impossible for scenes with differing geometry, we aim to match the local motion of the ground in the vicinity of the subject. This is achieved by reordering and warping the available target background frames in a manner that optimizes a suitably designed objective function. Fan Zhong 0001, Xueying Qin, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2013 | Deblurring by Example Using Dense CorrespondenceabstractThis paper presents a new method for deblurring photos using a sharp reference example that contains some shared content with the blurry photo. Most previous deblurring methods that exploit information from other photos require an accurately registered photo of the same static scene. In contrast, our method aims to exploit reference images where the shared content may have undergone substantial photometric and non-rigid geometric transformations, as these are the kind of reference images most likely to be found in personal photo albums. Our approach builds upon a recent method for example-based deblurring using non-rigid dense correspondence (NRDC) [HaCohen et al. 2011] and extends it in two ways. First, we suggest exploiting information from the reference image not only for blur kernel estimation, but also as a powerful local prior for the non-blind deconvolution step. Second, we introduce a simple yet robust technique for spatially varying blur estimation, rather than assuming spatially uniform blur. Unlike the above previous method, which has proven successful only with simple deblurring scenarios, we demonstrate that our method succeeds on a variety of real-world examples. We provide quantitative and qualitative evaluation of our method and show that it outperforms the state-of-the-art. Yoav HaCohen, Eli Shechtman, Dani Lischinski |
ICCV | 3 |
| 2013 | Illuminant Chromaticity from Image SequencesabstractWe estimate illuminant chromaticity from temporal sequences, for scenes illuminated by either one or two dominant illuminants. While there are many methods for illuminant estimation from a single image, few works so far have focused on videos, and even fewer on multiple light sources. Our aim is to leverage information provided by the temporal acquisition, where either the objects or the camera or the light source are/is in motion in order to estimate illuminant color without the need for user interaction or using strong assumptions and heuristics. We introduce a simple physically-based formulation based on the assumption that the incident light chromaticity is constant over a short space-time domain. We show that a deterministic approach is not sufficient for accurate and robust estimation: however, a probabilistic formulation makes it possible to implicitly integrate away hidden factors that have been ignored by the physical model. Experimental results are reported on a dataset of natural video sequences and on the Gray Ball benchmark, indicating that we compare favorably with the state-of-the-art. Véronique Prinet, Dani Lischinski, Michael Werman |
ICCV | 2 |
| 2013 | Specular highlight enhancement from video sequencesabstractWe propose a novel method for detecting and enhancing specular highlights from a video sequence, thereby obtaining a better visual perception of the specularity. To this end, we first generate a specularity map, defined over the space-time domain, by leveraging information provided by the temporal acquisition. We then amplify the highlights in each video frame to create the visual sensation of high dynamic range data. Results are illustrated on several videos taken under different acquisition scenarios. Véronique Prinet, Michael Werman, Dani Lischinski |
ICIP | 3 |
| 2013 | Locally Adaptive Products for All-Frequency RelightingabstractAbstract Triple product integrals evaluate the shading at a point by factoring the reflection equation into incident illumination, visibility, and BRDF. By densely sampling the space of incident directions, this approach is capable of highly accurate rendering scenes lit by high‐frequency environment lighting, containing complex materials and featuring intricate shadows. Efficient evaluation of triple product integrals using Haar wavelets enables near‐interactive rendering of such scenes, while dynamically changing the lighting and the view. Although faster methods have been proposed in the recent real‐time rendering literature, the approximations employed in these methods typically limit them to lower frequency phenomena. In this paper, we present a new approach for high‐frequency scene relighting within the triple product framework. Our approach breaks the computation to smaller solid angles (blocks) over most of which the triple product degenerates to a dot product. We introduce a lossless, yet compact, differential representation of the visibility function over each block, and sample the BRDF on the fly, eliminating the need to store multiple rotated copies of each BRDF. By combining these ideas, we are able to achieve true interactive performance even when running on a CPU, while supporting high frequency effects in scenes with high vertex counts. Yaron Inger, Zeev Farbman, Dani Lischinski |
Comput. Graph. Forum | 3 |
| 2013 | "Mind the gap": tele-registration for structure-driven image completionabstractConcocting a plausible composition from several non-overlapping image pieces, whose relative positions are not fixed in advance and without having the benefit of priors, can be a daunting task. Here we propose such a method, starting with a set of sloppily pasted image pieces with gaps between them. We first extract salient curves that approach the gaps from non-tangential directions, and use likely correspondences between pairs of such curves to guide a novel tele-registration method that simultaneously aligns all the pieces together. A structure-driven image completion technique is then proposed to fill the gaps, allowing the subsequent employment of standard in-painting tools to finish the job. Hui Huang 0004, Kangxue Yin, Minglun Gong, Dani Lischinski, Daniel Cohen-Or, Uri M. Ascher, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2013 | Optimizing color consistency in photo collectionsabstractWith dozens or even hundreds of photos in today's digital photo albums, editing an entire album can be a daunting task. Existing automatic tools operate on individual photos without ensuring consistency of appearance between photographs that share content. In this paper, we present a new method for consistent editing of photo collections. Our method automatically enforces consistent appearance of images that share content without any user input. When the user does make changes to selected images, these changes automatically propagate to other images in the collection, while still maintaining as much consistency as possible. This makes it possible to interactively adjust an entire photo album in a consistent manner by manipulating only a few images. Our method operates by efficiently constructing a graph with edges linking photo pairs that share content. Consistent appearance of connected photos is achieved by globally optimizing a quadratic cost function over the entire graph, treating user-specified edits as constraints in the optimization. The optimization is fast enough to provide interactive visual feedback to the user. We demonstrate the usefulness of our approach using a number of personal and professional photo collections, as well as internet collections. Yoav HaCohen, Eli Shechtman, Dan B. Goldman, Dani Lischinski |
ACM Trans. Graph. | 4 |
| 2012 | Content-Aware Automatic Photo EnhancementabstractAbstract Automatic photo enhancement is one of the long‐standing goals in image processing and computational photography. While a variety of methods have been proposed for manipulating tone and colour, most automatic methods used in practice, operate on the entire image without attempting to take the content of the image into account. In this paper, we present a new framework for automatic photo enhancement that attempts to take local and global image semantics into account. Specifically, our content‐aware scheme attempts to detect and enhance the appearance of human faces, blue skies with or without clouds and underexposed salient regions. A user study was conducted that demonstrates the effectiveness of the proposed approach compared to existing auto‐enhancement tools. Liad Kaufman, Dani Lischinski, Michael Werman |
Comput. Graph. Forum | 2 |
| 2012 | Digital reconstruction of halftoned color comicsabstractWe introduce a method for automated conversion of scanned color comic books and graphical novels into a new high-fidelity rescalable digital representation. Since crisp black line artwork and lettering are the most important structural and stylistic elements in this important genre of color illustrations, our digitization process is geared towards faithful reconstruction of these elements. This is a challenging task, because commercial presses perform halftoning (screening) to approximate continuous tones and colors with overlapping grids of dots. Although a large number of inverse haftoning (descreening) methods exist, they typically blur the intricate black artwork. Our approach is specifically designed to descreen color comics, which typically reproduce color using screened CMY inks, but print the black artwork using non-screened solid black ink. After separating the scanned image into three screening grids, one for each of the CMY process inks, we use non-linear optimization to fit a parametric model describing each grid, and simultaneously recover the non-screened black ink layer, which is then vectorized. The result of this process is a high quality, compact, and rescalable digital representation of the original artwork. Johannes Kopf 0001, Dani Lischinski |
ACM Trans. Graph. | 2 |
| 2011 | On Neighbourhood Matching for Texture-by-NumbersabstractAbstract Texture‐by‐Numbers is an attractive texture synthesis framework, because it is able to cope with non‐homogeneous texture exemplars, and provides the user with intuitive creative control over the outcome of the synthesis process. Like many other exemplar‐based texture synthesis methods, its basic underlying mechanism is neighbourhood matching. In this paper we review a number of commonly used neighbourhood matching acceleration techniques, compare and analyse their performance in the specific context of Texture‐by‐Numbers (as opposed to ordinary unconstrained texture synthesis). Our study indicates that the standard approaches are not optimally suited for the Texture‐by‐Numbers framework, often producing visually inferior results compared to searching for the exact L2nearest neighbour. We then show that performing Texture‐by‐Number using the Texture Optimization framework in conjunction with an efficient FFT‐based search is able to produce good results in reasonable running times and with a minimal memory overhead. Eliyahu Sivaks, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2011 | Convolution pyramidsabstractWe present a novel approach for rapid numerical approximation of convolutions with filters of large support. Our approach consists of a multiscale scheme, fashioned after the wavelet transform, which computes the approximation in linear time. Given a specific large target filter to approximate, we first use numerical optimization to design a set of small kernels, which are then used to perform the analysis and synthesis steps of our multiscale transform. Once the optimization has been done, the resulting transform can be applied to any signal in linear time. We demonstrate that our method is well suited for tasks such as gradient field integration, seamless image cloning, and scattered data interpolation, outperforming existing state-of-the-art methods. Zeev Farbman, Raanan Fattal, Dani Lischinski |
ACM Trans. Graph. | 3 |
| 2011 | Tonal stabilization of videoabstractThis paper presents a method for reducing undesirable tonal fluctuations in video: minute changes in tonal characteristics, such as exposure, color temperature, brightness and contrast in a sequence of frames, which are easily noticeable when the sequence is viewed. These fluctuations are typically caused by the camera's automatic adjustment of its tonal settings while shooting. Our approach operates on a continuous video shot by first designating one or more frames as anchors . We then tonally align a sequence of frames with each anchor: for each frame, we compute an adjustment map that indicates how each of its pixels should be modified in order to appear as if it was captured with the tonal settings of the anchor. The adjustment map is efficiently updated between successive frames by taking advantage of temporal video coherence and the global nature of the tonal fluctuations. Once a sequence has been aligned, it is possible to generate smooth tonal transitions between anchors, and also further control its tonal characteristics in a consistent and principled manner, which is difficult to do without incurring strong artifacts when operating on unstable sequences. We demonstrate the utility of our method using a number of clips captured with a variety of video cameras, and believe that it is well-suited for integration into today's non-linear video editing tools. Zeev Farbman, Dani Lischinski |
ACM Trans. Graph. | 2 |
| 2011 | Non-rigid dense correspondence with applications for image enhancementabstractThis paper presents a new efficient method for recovering reliable local sets of dense correspondences between two images with some shared content. Our method is designed for pairs of images depicting similar regions acquired by different cameras and lenses, under non-rigid transformations, under different lighting, and over different backgrounds. We utilize a new coarse-to-fine scheme in which nearest-neighbor field computations using Generalized PatchMatch [Barnes et al. 2010] are interleaved with fitting a global non-linear parametric color model and aggregating consistent matching regions using locally adaptive constraints. Compared to previous correspondence approaches, our method combines the best of two worlds: It is dense, like optical flow and stereo reconstruction methods, and it is also robust to geometric and photometric variations, like sparse feature matching. We demonstrate the usefulness of our method using three applications for automatic example-based photograph enhancement: adjusting the tonal characteristics of a source image to match a reference, transferring a known mask to a new image, and kernel estimation for image deblurring. Yoav HaCohen, Eli Shechtman, Dan B. Goldman, Dani Lischinski |
ACM Trans. Graph. | 4 |
| 2011 | Depixelizing pixel artabstractWe describe a novel algorithm for extracting a resolution-independent vector representation from pixel art images, which enables magnifying the results by an arbitrary amount without image degradation. Our algorithm resolves pixel-scale features in the input and converts them into regions with smoothly varying shading that are crisply separated by piecewise-smooth contour curves. In the original image, pixels are represented on a square pixel lattice, where diagonal neighbors are only connected through a single point. This causes thin features to become visually disconnected under magnification by conventional means, and creates ambiguities in the connectedness and separation of diagonal neighbors. The key to our algorithm is in resolving these ambiguities. This enables us to reshape the pixel cells so that neighboring pixels belonging to the same feature are connected through edges, thereby preserving the feature connectivity under magnification. We reduce pixel aliasing artifacts and improve smoothness by fitting spline curves to contours in the image and optimizing their control points. Johannes Kopf 0001, Dani Lischinski |
ACM Trans. Graph. | 2 |
| 2010 | Personalization of image enhancementabstractWe address the problem of incorporating user preference in automatic image enhancement. Unlike generic tools for automatically enhancing images, we seek to develop methods that can first observe user preferences on a training set, and then learn a model of these preferences to personalize enhancement of unseen images. The challenge of designing such system lies at intersection of computer vision, learning, and usability; we use techniques such as active sensor selection and distance metric learning in order to solve the problem. The experimental evaluation based on user studies indicates that different users do have different preferences in image enhancement, which suggests that personalization can further help improve the subjective quality of generic image enhancements. Sing Bing Kang, Ashish Kapoor, Dani Lischinski |
CVPR | 3 |
| 2010 | Diffusion maps for edge-aware image editingabstractEdge-aware operations, such as edge-preserving smoothing and edge-aware interpolation, require assessing the degree of similarity between pairs of pixels, typically defined as a simple monotonic function of the Euclidean distance between pixel values in some feature space. In this work we introduce the idea of replacing these Euclidean distances with diffusion distances , which better account for the global distribution of pixels in their feature space. These distances are approximated using diffusion maps : a set of the dominant eigenvectors of a large affinity matrix, which may be computed efficiently by sampling a small number of matrix columns (the Nyström method). We demonstrate the benefits of using diffusion distances in a variety of image editing contexts, and explore the use of diffusion maps as a tool for facilitating the creation of complex selection masks. Finally, we present a new analysis that establishes a connection between the spatial interaction range between two pixels, and the number of samples necessary for accurate Nyström approximations. Zeev Farbman, Raanan Fattal, Dani Lischinski |
ACM Trans. Graph. | 3 |
| 2009 | Locally Adapted Projections to Reduce Panorama DistortionsabstractAbstract Displaying panoramic and wide angle views on a flat 2D display surface is necessarily prone to distortions. Perspective projections are limited to fairly narrow view angles. Cylindrical and spherical projections can show full 360° panoramas, but at the cost of curving straight lines, interfering with the perception of salient shapes in the scene. In this paper, we introducelocally‐adapted projections. Such projections are defined by a continuous projection surface consisting of both near‐planar and curved parts. A simple and intuitive user interface allows the specification of regions of interest to be mapped to the near‐planar parts, thereby reducing bending artifacts. We demonstrate the effectiveness of our approach on a variety of panoramic and wide angle images, including both indoor and outdoor scenes. Johannes Kopf 0001, Dani Lischinski, Oliver Deussen, Daniel Cohen-Or, Michael F. Cohen |
Comput. Graph. Forum | 2 |
| 2009 | Coordinates for instant image cloningabstractSeamless cloning of a source image patch into a target image is an important and useful image editing operation, which has received considerable research attention in recent years. This operation is typically carried out by solving a Poisson equation with Dirichlet boundary conditions, which smoothly interpolates the discrepancies between the boundary of the source patch and the target across the entire cloned area. In this paper we introduce an alternative, coordinate-based approach, where rather than solving a large linear system to perform the aforementioned interpolation, the value of the interpolant at each interior pixel is given by a weighted combination of values along the boundary. More specifically, our approach is based on Mean-Value Coordinates (MVC). The use of coordinates is advantageous in terms of speed, ease of implementation, small memory footprint, and parallelizability, enabling real-time cloning of large regions, and interactive cloning of video streams. We demonstrate a number of applications and extensions of the coordinate-based framework. Zeev Farbman, Gil Hoffer, Yaron Lipman, Daniel Cohen-Or, Dani Lischinski |
ACM Trans. Graph. | 5 |
| 2009 | Layered shape synthesis: automatic generation of control maps for non-stationary texturesabstractMany inhomogeneous real-world textures are non-stationary and exhibit various large scale patterns that are easily perceived by a human observer. Such textures violate the assumptions underlying most state-of-the-art example-based synthesis methods. Consequently, they cannot be properly reproduced by these methods, unless a suitable control map is provided to guide the synthesis process. Such control maps are typically either user specified or generated by a simulation. In this paper, we present an alternative: a method for automatic example-based generation of control maps, geared at synthesis of natural, highly inhomogeneous textures, such as those resulting from natural aging or weathering processes. Our method is based on the observation that an appropriate control map for many of these textures may be modeled as a superposition of several layers, where the visible parts of each layer are occupied by a more homogeneous texture. Thus, given a decomposition of a texture exemplar into a small number of such layers, we employ a novel example-based shape synthesis algorithm to automatically generate a new set of layers. Our shape synthesis algorithm is designed to preserve both local and global characteristics of the exemplar's layer map. This process results in a new control map, which then may be used to guide the subsequent texture synthesis process. Amir Rosenberger, Daniel Cohen-Or, Dani Lischinski |
ACM Trans. Graph. | 3 |
| 2008 | The Shadow Meets the Mask: Pyramid-Based Shadow RemovalabstractAbstract In this paper we propose a novel method for detecting and removing shadows from a single image thereby obtaining a high‐quality shadow‐free image. With minimal user assistance, we first identify shadowed and lit areas on the same surface in the scene using an illumination‐invariant distance measure. These areas are used to estimate the parameters of an affine shadow formation model. A novel pyramid‐based restoration process is then applied to produce a shadow‐free image, while avoiding loss of texture contrast and introduction of noise. Unlike previous approaches, we account for varying shadow intensity inside the shadowed region by processing it from the interior towards the boundaries. Finally, to ensure a seamless transition between the original and the recovered regions we apply image inpainting along a thin border. We demonstrate that our approach produces results that are in most cases superior in quality to those of previous shadow removal methods. We also show that it is possible to easily composite the extracted shadow onto a new background or modify its size and direction in the original image. Yael Shor, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2008 | A Closed-Form Solution to Natural Image MattingabstractInteractive digital matting, the process of extracting a foreground object from an image based on limited user input, is an important task in image and video editing. From a computer vision perspective, this task is extremely challenging because it is massively ill-posed -- at each pixel we must estimate the foreground and the background colors, as well as the foreground opacity ("alpha matte") from a single color measurement. Current approaches either restrict the estimation to a small part of the image, estimating foreground and background colors based on nearby pixels where they are known, or perform iterative nonlinear estimation by alternating foreground and background color estimation with alpha estimation. In this paper we present a closed-form solution to natural image matting. We derive a cost function from local smoothness assumptions on foreground and background colors, and show that in the resulting expression it is possible to analytically eliminate the foreground and background colors to obtain a quadratic cost function in alpha. This allows us to find the globally optimal alpha matte by solving a sparse linear system of equations. Furthermore, the closed-form formula allows us to predict the properties of the solution by analyzing the eigenvectors of a sparse matrix, closely related to matrices used in spectral image segmentation algorithms. We show that high quality mattes for natural images may be obtained from a small amount of user input. Anat Levin, Dani Lischinski, Yair Weiss |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Spectral MattingabstractWe present spectral matting: a new approach to natural image matting that automatically computes a basis set of fuzzy matting components from the smallest eigenvectors of a suitably defined Laplacian matrix. Thus, our approach extends spectral segmentation techniques, whose goal is to extract hard segments, to the extraction of soft matting components. These components may then be used as building blocks to easily construct semantically meaningful foreground mattes, either in an unsupervised fashion, or based on a small amount of user input. Anat Levin, Alex Rav-Acha, Dani Lischinski |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Edge-preserving decompositions for multi-scale tone and detail manipulationabstractMany recent computational photography techniques decompose an image into a piecewise smooth base layer, containing large scale variations in intensity, and a residual detail layer capturing the smaller scale details in the image. In many of these applications, it is important to control the spatial scale of the extracted details, and it is often desirable to manipulate details at multiple scales, while avoiding visual artifacts. In this paper we introduce a new way to construct edge-preserving multi-scale image decompositions. We show that current basedetail decomposition techniques, based on the bilateral filter, are limited in their ability to extract detail at arbitrary scales. Instead, we advocate the use of an alternative edge-preserving smoothing operator, based on the weighted least squares optimization framework, which is particularly well suited for progressive coarsening of images and for multi-scale detail extraction. After describing this operator, we show how to use it to construct edge-preserving multi-scale decompositions, and compare it to the bilateral filter, as well as to other schemes. Finally, we demonstrate the effectiveness of our edge-preserving decompositions in the context of LDR and HDR tone mapping, detail enhancement, and other applications. Zeev Farbman, Raanan Fattal, Dani Lischinski, Richard Szeliski |
ACM Trans. Graph. | 3 |
| 2008 | Deep photo: model-based photograph enhancement and viewingabstractIn this paper, we introduce a novel system for browsing, enhancing, and manipulating casual outdoor photographs by combining them with already existing georeferenced digital terrain and urban models. A simple interactive registration process is used to align a photograph with such a model. Once the photograph and the model have been registered, an abundance of information, such as depth, texture, and GIS data, becomes immediately available to our system. This information, in turn, enables a variety of operations, ranging from dehazing and relighting the photograph, to novel view synthesis, and overlaying with geographic information. We describe the implementation of a number of these applications and discuss possible extensions. Our results show that augmenting photographs with already available 3D models of the world supports a wide variety of new ways for us to experience and interact with our everyday snapshots. Johannes Kopf 0001, Boris Neubert, Billy Chen, Michael F. Cohen, Daniel Cohen-Or, Oliver Deussen, Matthew Uyttendaele, Dani Lischinski |
ACM Trans. Graph. | 8 |
| 2008 | Data-driven enhancement of facial attractivenessabstractWhen human raters are presented with a collection of shapes and asked to rank them according to their aesthetic appeal, the results often indicate that there is a statistical consensus among the raters. Yet it might be difficult to define a succinct set of rules that capture the aesthetic preferences of the raters. In this work, we explore a data-driven approach to aesthetic enhancement of such shapes. Specifically, we focus on the challenging problem of enhancing the aesthetic appeal (or the attractiveness ) of human faces in frontal photographs (portraits), while maintaining close similarity with the original. The key component in our approach is an automatic facial attractiveness engine trained on datasets of faces with accompanying facial attractiveness ratings collected from groups of human raters. Given a new face, we extract a set of distances between a variety of facial feature locations, which define a point in a high-dimensional "face space". We then search the face space for a nearby point with a higher predicted attractiveness rating. Once such a point is found, the corresponding facial distances are embedded in the plane and serve as a target to define a 2D warp field which maps the original facial features to their adjusted locations. The effectiveness of our technique was experimentally validated by independent rating experiments, which indicate that it is indeed capable of increasing the facial attractiveness of most portraits that we have experimented with. Tommer Leyvand, Daniel Cohen-Or, Gideon Dror, Dani Lischinski |
ACM Trans. Graph. | 4 |
| 2007 | Spectral MattingabstractWe present spectral matting: a new approach to natural image matting that automatically computes a set of fundamental fuzzy matting components from the smallest eigenvectors of a suitably defined Laplacian matrix. Thus, our approach extends spectral segmentation techniques, whose goal is to extract hard segments, to the extraction of soft matting components. These components may then be used as building blocks to easily construct semantically meaningful foreground mattes, either in an unsupervised fashion, or based on a small amount of user input. Anat Levin, Alex Rav-Acha, Dani Lischinski |
CVPR | 3 |
| 2007 | Crowds by ExampleabstractAbstract We present an example‐based crowd simulation technique. Most crowd simulation techniques assume that the behavior exhibited by each person in the crowd can be defined by a restricted set of rules. This assumption limits the behavioral complexity of the simulated agents. By learning from real‐world examples, our autonomous agents display complex natural behaviors that are often missing in crowd simulations. Examples are created from tracked video segments of real pedestrian crowds. During a simulation, autonomous agents search for examples that closely match the situation that they are facing. Trajectories taken by real people in similar situations, are copied to the simulated agents, resulting in seemingly natural behaviors. Alon Lerner, Yiorgos Chrysanthou, Dani Lischinski |
Comput. Graph. Forum | 3 |
| 2007 | Dynamosaicing: Mosaicing of Dynamic ScenesabstractThis paper explores the manipulation of time in video editing, enabling to control the chronological time of events. These time manipulations include slowing down (or postponing) some dynamic events while speeding up (or advancing) others. When a video camera scans a scene, aligning all the events to a single time interval will result in a panoramic movie. Time manipulations are obtained by first constructing an aligned space-time volume from the input video, and then sweeping a continuous 2D slice (time front) through that volume, generating a new sequence of images. For dynamic scenes, aligning the input video frames poses an important challenge. We propose to align dynamic scenes using a new notion of "dynamics constancy", which is more appropriate for this task than the traditional assumption of "brightness constancy". Another challenge is to avoid visual seams inside moving objects and other visual artifacts resulting from sweeping the space-time volumes with time fronts of arbitrary geometry. To avoid such artifacts, we formulate the problem of finding optimal time front geometry as one of finding a minimal cut in a 4D graph, and solve it using max-flow methods. Alex Rav-Acha, Yael Pritch, Dani Lischinski, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Joint bilateral upsamplingabstractImage analysis and enhancement tasks such as tone mapping, colorization, stereo depth, and photomontage, often require computing a solution (e.g., for exposure, chromaticity, disparity, labels) over the pixel grid. Computational and memory costs often require that a smaller solution be run over a downsampled image. Although general purpose upsampling methods can be used to interpolate the low resolution solution to the full resolution, these methods generally assume a smoothness prior for the interpolation. We demonstrate that in cases, such as those above, the available high resolution input image may be leveraged as a prior in the context of a joint bilateral upsampling procedure to produce a better high resolution solution. We show results for each of the applications above and compare them to traditional upsampling methods. Johannes Kopf 0001, Michael F. Cohen, Dani Lischinski, Matthew Uyttendaele |
ACM Trans. Graph. | 3 |
| 2007 | Solid texture synthesis from 2D exemplarsabstractWe present a novel method for synthesizing solid textures from 2D texture exemplars. First, we extend 2D texture optimization techniques to synthesize 3D texture solids. Next, the non-parametric texture optimization approach is integrated with histogram matching, which forces the global statistics of the synthesized solid to match those of the exemplar. This improves the convergence of the synthesis process and enables using smaller neighborhoods. In addition to producing compelling texture mapped surfaces, our method also effectively models the material in the interior of solid objects. We also demonstrate that our method is well-suited for synthesizing textures with a large number of channels per texel. Johannes Kopf 0001, Chi-Wing Fu, Daniel Cohen-Or, Oliver Deussen, Dani Lischinski, Tien-Tsin Wong |
ACM Trans. Graph. | 5 |
| 2006 | Inducing Semantic Segmentation from an Example
Yaar Schnitman, Yaron Caspi, Daniel Cohen-Or, Dani Lischinski |
ACCV (2) | 4 |
| 2006 | A Closed Form Solution to Natural Image MattingabstractInteractive digital matting, the process of extracting a foreground object from an image based on limited user input, is an important task in image and video editing. From a computer vision perspective, this task is extremely challenging because it is massively ill-posed - at each pixel we must estimate the foreground and the background colors, as well as the foreground opacity ("alpha matte") from a single color measurement. Current approaches either restrict the estimation to a small part of the image, estimating foreground and background colors based on nearby pixels where they are known, or perform iterative nonlinear estimation by alternating foreground and background color estimation with alpha estimation. In this paper we present a closed form solution to natural image matting. We derive a cost function from local smoothness assumptions on foreground and background colors, and show that in the resulting expression it is possible to analytically eliminate the foreground and background colors to obtain a quadratic cost function in alpha. This allows us to find the globally optimal alpha matte by solving a sparse linear system of equations. Furthermore, the closed form formula allows us to predict the properties of the solution by analyzing the eigenvectors of a sparse matrix, closely related to matrices used in spectral image segmentation algorithms. We show that high quality mattes can be obtained on natural images from a surprisingly small amount of user input. Anat Levin, Dani Lischinski, Yair Weiss |
CVPR (1) | 2 |
| 2006 | Pose Controlled Physically Based MotionabstractAbstract In this paper we describe a new method for generating and controlling physically‐based motion of complex articulated characters. Our goal is to create motion from scratch, where the animator provides a small amount of input and gets in return a highly detailed and physically plausible motion. Our method relieves the animator from the burden of enforcing physical plausibility, but at the same time provides full control over the internal DOFs of the articulated character via a familiar interface. Control over the global DOFs is also provided by supporting kinematic constraints. Unconstrained portions of the motion are generated in real time, since the character is driven by joint torques generated by simple feedback controllers. Although kinematic constraints are satisfied using an iterative search (shooting), this process is typically inexpensive, since it only adjusts a few DOFs at a few time instances. The low expense of the optimization, combined with the ability to generate unconstrained motions in real time yields an efficient and practical tool, which is particularly attractive for high inertia motions with a relatively small number of kinematic constraints. Raanan Fattal, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2006 | Recursive Wang tiles for real-time blue noiseabstractWell distributed point sets play an important role in a variety of computer graphics contexts, such as anti-aliasing, global illumination, halftoning, non-photorealistic rendering, point-based modeling and rendering, and geometry processing. In this paper, we introduce a novel technique for rapidly generating large point sets possessing a blue noise Fourier spectrum and high visual quality. Our technique generates non-periodic point sets, distributed over arbitrarily large areas. The local density of a point set may be prescribed by an arbitrary target density function, without any preset bound on the maximum density. Our technique is deterministic and tile-based; thus, any local portion of a potentially infinite point set may be consistently regenerated as needed. The memory footprint of the technique is constant, and the cost to generate any local portion of the point set is proportional to the integral over the target density in that area. These properties make our technique highly suitable for a variety of real-time interactive applications, some of which are demonstrated in the paper.Our technique utilizes a set of carefully constructed progressive and recursive blue noise Wang tiles. The use of Wang tiles enables the generation of infinite non-periodic tilings. The progressive point sets inside each tile are able to produce spatially varying point densities. Recursion allows our technique to adaptively subdivide tiles only where high density is required, and makes it possible to zoom into point sets by an arbitrary amount, while maintaining a constant apparent density. Johannes Kopf 0001, Daniel Cohen-Or, Oliver Deussen, Dani Lischinski |
ACM Trans. Graph. | 4 |
| 2006 | Interactive local adjustment of tonal valuesabstractThis paper presents a new interactive tool for making local adjustments of tonal values and other visual parameters in an image. Rather than carefully selecting regions or hand-painting layer masks, the user quickly indicates regions of interest by drawing a few simple brush strokes and then uses sliders to adjust the brightness, contrast, and other parameters in these regions. The effects of the user's sparse set of constraints are interpolated to the entire image using an edge-preserving energy minimization method designed to prevent the propagation of tonal adjustments to regions of significantly different luminance. The resulting system is suitable for adjusting ordinary and high dynamic range images, and provides the user with much more creative control than existing tone mapping algorithms. Our tool is also able to produce a tone mapping automatically, which may serve as a basis for further local adjustments, if so desired. The constraint propagation approach developed in this paper is a general one, and may also be used to interactively control a variety of other adjustments commonly performed in the digital darkroom. Dani Lischinski, Zeev Farbman, Matthew Uyttendaele, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2006 | Guest editorial
Deok-Soo Kim, In-Kwon Lee, Dani Lischinski, Ayellet Tal |
Vis. Comput. | 3 |
| 2005 | Dynamosaics: Video Mosaics with Non-Chronological TimeabstractWith the limited field of view of human vision, our perception of most scenes is built over time while our eyes are scanning the scene. In the case of static scenes, this process can be modeled by panoramic mosaicing: stitching together images into a panoramic view. Can a dynamic scene, scanned by a video camera, be represented with a dynamic panoramic video even though different regions were visible at different times? In this paper, we explore time flow manipulation in video, such as the creation of new videos in which events that occurred at different times are displayed simultaneously. More general changes in the time flow are also possible, which enable re-scheduling the order of dynamic events in the video, for example. We generate dynamic mosaics by sweeping the aligned space-time volume of the input video by a time front surface and generating a sequence of time slices in the process. Various sweeping strategies and different time front evolutions manipulate the time flow in the video, enabling many unexplored and powerful effects, such as panoramic movies. Alex Rav-Acha, Yael Pritch, Dani Lischinski, Shmuel Peleg |
CVPR (1) | 3 |
| 2005 | Colorization by Example
Revital Ironi, Daniel Cohen-Or, Dani Lischinski |
Rendering Techniques | 3 |
| 2004 | Target-driven smoke animationabstractIn this paper we present a new method for efficiently controlling animated smoke. Given a sequence of target smoke states, our method generates a smoke simulation in which the smoke is driven towards each of these targets in turn, while exhibiting natural-looking interesting smoke-like behavior. This control is made possible by two new terms that we add to the standard flow equations: (i) a driving force term that causes the fluid to carry the smoke towards a particular target, and (ii) a smoke gathering term that prevents the smoke from diffusing too much. These terms are explicitly defined by the instantaneous state of the system at each simulation timestep. Thus, no expensive optimization is required, allowing complex smoke animations to be generated with very little additional cost compared to ordinary flow simulations. Raanan Fattal, Dani Lischinski |
ACM Trans. Graph. | 2 |
| 2004 | Colorization using optimizationabstractColorization is a computer-assisted process of adding color to a monochrome image or movie. The process typically involves segmenting images into regions and tracking these regions across image sequences. Neither of these tasks can be performed reliably in practice; consequently, colorization requires considerable user intervention and remains a tedious, time-consuming, and expensive task.In this paper we present a simple colorization method that requires neither precise image segmentation, nor accurate region tracking. Our method is based on a simple premise; neighboring pixels in space-time that have similar intensities should have similar colors. We formalize this premise using a quadratic cost function and obtain an optimization problem that can be solved efficiently using standard techniques. In our approach an artist only needs to annotate the image with a few color scribbles, and the indicated colors are automatically propagated in both space and time to produce a fully colorized image or sequence. We demonstrate that high quality colorizations of stills and movie clips may be obtained from a relatively modest amount of user input. Anat Levin, Dani Lischinski, Yair Weiss |
ACM Trans. Graph. | 2 |
| 2004 | Constrained synthesis of textural motion for articulated characters
Shmuel Moradoff, Dani Lischinski |
Vis. Comput. | 2 |
| 2003 | Fast Multiresolution Image Operations in the Wavelet DomainabstractA wide class of operations on images can be performed directly in the wavelet domain by operating on coefficients of the wavelet transforms of the images and other matrices defined by these operations. Operating in the wavelet domain enables one to perform these operations progressively in a coarse-to-fine fashion, operate on different resolutions, manipulate features at different scales, trade off accuracy for speed, and localize the operation in both the spatial and the frequency domains. Performing such operations in the wavelet domain and then reconstructing the result is also often more efficient than performing the same operation in the standard direct fashion. In this paper, we demonstrate the applicability and advantages of this framework to three common types of image operations: image blending, 3D warping of images and sequences, and convolution of images and image sequences. Iddo Drori, Dani Lischinski |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2002 | Bounded-distortion Piecewise Mesh ParameterizationabstractMany computer graphics operations, such as texture mapping, 3D painting, remeshing, mesh compression, and digital geometry processing, require finding a low-distortion parameterization for irregular connectivity triangulations of arbitrary genus 2-manifolds. This paper presents a simple and fast method for computing parameterizations with strictly bounded distortion. The new method operates by flattening the mesh onto a region of the 2D plane. To comply with the distortion bound, the mesh is automatically cut and partitioned on-the-fly. The method guarantees avoiding global and local self-intersections, while attempting to minimize the total length of the introduced seams. To our knowledge, this is the first method to compute the mesh partitioning and the parameterization simultaneously and entirely automatically, while providing guaranteed distortion bounds. Our results on a variety of objects demonstrate that the method is fast enough to work with large complex irregular meshes in interactive applications. Olga Sorkine-Hornung, Daniel Cohen-Or, Rony Goldenthal, Dani Lischinski |
IEEE Visualization | 4 |
| 2002 | Gradient domain high dynamic range compressionabstractWe present a new method for rendering high dynamic range images on conventional displays. Our method is conceptually simple, computationally efficient, robust, and easy to use. We manipulate the gradient field of the luminance image by attenuating the magnitudes of large gradients. A new, low dynamic range image is then obtained by solving a Poisson equation on the modified gradient field. Our results demonstrate that the method is capable of drastic dynamic range compression, while preserving fine details and avoiding common artifacts, such as halos, gradient reversals, or loss of local contrast. The method is also able to significantly enhance ordinary images by bringing out detail in dark regions. Raanan Fattal, Dani Lischinski, Michael Werman |
ACM Trans. Graph. | 2 |
| 2001 | Variational Classification for Visualization of 3D Ultrasound DataabstractWe present a new technique for visualizing surfaces from 3D ultrasound data. 3D ultrasound datasets are typically fuzzy, contain a substantial amount of noise and speckle, and suffer from several other problems that make extraction of continuous and smooth surfaces extremely difficult. We propose a novel opacity classification algorithm for 3D ultrasound datasets, based on the variational principle. More specifically, we compute a volumetric opacity function that optimally satisfies a set of simultaneous requirements. One requirement makes the function attain nonzero values only in the vicinity of a user-specified value, resulting in soft shells of finite, approximately constant thickness around isosurfaces in the volume. Other requirements are designed to make the function smoother and less sensitive to noise and speckle. The computed opacity function lends itself well to explicit geometric surface extraction, as well as to direct volume rendering at interactive rates. We also describe a new splatting algorithm that is particularly well suited for displaying soft opacity shells. Several examples and comparisons are included to illustrate our approach and demonstrate its effectiveness on real 3D ultrasound datasets. Raanan Fattal, Dani Lischinski |
IEEE Visualization | 2 |
| 2001 | Automatic Lighting Design using a Perceptual Quality MetricabstractLighting has a crucial impact on the appearance of 3D objects and on the ability of an image to communicate information about a 3D scene to a human observer. This paper presents a new automatic lighting design approach for comprehensible rendering of 3D objects. Given a geometric model of a 3D object or scene, the material properties of the surfaces in the model, and the desired viewing parameters, our approach automatically determines the values of various lighting parameters by optimizing a perception-based image quality objective function. This objective function is designed to quantify the extent to which an image of a 3D scene succeeds in communicating scene information, such as the 3D shapes of the objects, fine geometric details, and the spatial relationships between the objects. Our results demonstrate that the proposed approach is an effective lighting design tool, suitable for users without expertise or knowledge in visual perception or in lighting design. Ram Shacked, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2001 | Streaming of Complex 3D Scenes for Remote WalkthroughsabstractWe describe a new 3D scene streaming approach for remote walkthroughs. In a remote walkthrough, a user on a client machine interactively navigates through a scene that resides on a remote server. Our approach allows a user to walk through a remote 3D scene, without ever having to download the entire scene from the server. Our algorithm achieves this by selectively transmitting only small parts of the scene and lower quality representations of objects, based on the user's viewing parameters and the available connection bandwidth. An online optimization algorithm selects which object representations to send, based on the integral of a benefit measure along the predicted path of movement. The rendering quality at the client depends on the available bandwidth, but practical navigation of the scene is possible even when bandwidth is low. Eyal Teler, Dani Lischinski |
Comput. Graph. Forum | 2 |
| 2001 | Texture Mixing and Texture Movie Synthesis Using Statistical LearningabstractWe present an algorithm based on statistical learning for synthesizing static and time-varying textures matching the appearance of an input texture. Our algorithm is general and automatic and it works well on various types of textures, including 1D sound textures, 2D texture images, and 3D texture movies. The same method is also used to generate 2D texture mixtures that simultaneously capture the appearance of a number of different input textures. In our approach, input textures are treated as sample signals generated by a stochastic process. We first construct a tree representing a hierarchical multiscale transform of the signal using wavelets. From this tree, new random trees are generated by learning and sampling the conditional probabilities of the paths in the original tree. Transformation of these random trees back into signals results in new random textures. In the case of 2D texture synthesis, our algorithm produces results that are generally as good as or better than those produced by previously described methods in this field. For texture mixtures, our results are better and more general than those produced by earlier methods. For texture movies, we present the first algorithm that is able to automatically generate movie clips of dynamic phenomena such as waterfalls, fire flames, a school of jellyfish, a crowd of people, etc. Our results indicate that the proposed technique is effective and robust. Ziv Bar-Joseph, Ran El-Yaniv, Dani Lischinski, Michael Werman |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2000 | Compression of Indoor Video Sequences Using Homography-Based SegmentationabstractWe present a new compression algorithm for video sequences of indoor scenes, or more generally, sequences containing mostly planar and near-planar surfaces. Our approach utilizes edge and optical flow information in order to segment selected keyframes into regions of general shape, such that the motion of each region is predicted well by a planar homography. With this kind of motion prediction, the errors between the predicted intermediate frames and the actual ones are very small, and can be compactly encoded. Our results demonstrate significant improvements in the accuracy of the compressed video sequences, compared to standard general purpose video compression. Tae-Joon Park, Shachar Fleishman, Daniel Cohen-Or, Dani Lischinski |
PG | 4 |
| 2000 | Automatic Camera Placement for Image-Based ModelingabstractWe present an automatic camera placement method for generating image‐based models from scenes with known geometry. Our method first approximately determines the set of surfaces visible from a given viewing area and then selects a small set of appropriate camera positions to sample the scene from. We define a quality measure for a surface as seen, or covered, from the given viewing area. Along with each camera position, we store the set of surfaces which are best covered by this camera. Next, one reference view is generated from each camera position by rendering the scene. Pixels in each reference view that do not belong to the selected set of polygons are masked out. The image‐based model generated by our method, covers every visible surface only once, associating it with a camera position from which it is covered with quality that exceeds a user‐specified quality threshold. The result is a compact non‐redundant image‐based model with controlled quality. The problem of covering every visible surface with a minimum number of cameras (guards) can be regarded as an extension to the well‐known Art Gallery Problem. However, since the 3D polygonal model is textured, the camera‐polygon visibility relation is not binary; instead, it has a weight — the quality of the polygon's coverage. Shachar Fleishman, Daniel Cohen-Or, Dani Lischinski |
Comput. Graph. Forum | 3 |
| 1999 | Automatic Camera Placement for Image-Based ModelingabstractWe present an automatic camera placement method for generating image-based models from scenes with known geometry. Our method first approximately determines the set of surfaces visible from a given viewing area and then selects a small set of appropriate camera positions to sample the scene from. We define a quality measure for a surface as seen, or covered, from the given viewing area. Along with each camera position, we store the set of surfaces which are best covered by this camera. Next, one reference view is generated from each reference view that do not belong to the selected set of polygons are masked out. The image-based model generated by our method, covers every visible surface only once, associating it with a camera position from which it is covered with quality that exceeds a user-specified quality threshold. The result is a compact non-redundant image-based model with controlled quality. The problem of covering every visible surface with a minimum number of cameras (guards) can be regarded as an extension to the well-known Art Gallery Problem. However, since the 3D polygonal model is textured, the camera-polygon visibility relation is not binary; instead, it has a weight-the quality of the polygon's coverage. Shachar Fleishman, Daniel Cohen-Or, Dani Lischinski |
PG | 3 |
| 1998 | Fast Approximate Quantitative Visibility for Complex ScenesabstractRay tracing and Monte-Carlo based global illumination, as well as radiosity and other finite-element based global illumination methods, all require repeated evaluation of quantitative visibility queries, such as: what is the average visibility between a point (a differential area element) and a finite area or volume; or what is the average visibility between two finite areas or volumes. We present a new data structure and an algorithm for rapidly evaluating such queries in complex scenes. The proposed approach utilizes a novel image-based discretization of the space of bounded rays in the scene, constructed in a preprocessing stage. This data structure makes it possible to quickly compute approximate answers to visibility queries. Because visibility queries are computed using a discretization of the space, the execution time is effectively decoupled from the number of geometric primitives in the scene. A potential hazard with the proposed approach is that it might require large amounts of memory, if the data structures are designed in a naive fashion. We discuss ways for representing the discretization in a compact manner while still allowing rapid query evaluation. Preliminary results demonstrate the effectiveness of the proposed approach. Yiorgos Chrysanthou, Daniel Cohen-Or, Dani Lischinski |
Computer Graphics International | 3 |
| 1998 | Synthesizing Realistic Facial Expressions from PhotographsabstractArticle Free Access Share on Synthesizing realistic facial expressions from photographs Authors: Frédéric Pighin Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile , Jamie Hecker Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile , Dani Lischinski Hebrew Univ. Hebrew Univ.View Profile , Richard Szeliski Microsoft Research Microsoft ResearchView Profile , David H. Salesin Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile Authors Info & Claims SIGGRAPH '98: Proceedings of the 25th annual conference on Computer graphics and interactive techniquesJuly 1998 Pages 75–84https://doi.org/10.1145/280814.280825Published:24 July 1998Publication History 446citation2,172DownloadsMetricsTotal Citations446Total Downloads2,172Last 12 Months71Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Frédéric H. Pighin, Jamie Hecker, Dani Lischinski, Richard Szeliski, David Salesin |
SIGGRAPH | 3 |
| 1997 | Clustering for Glossy Global IlluminationabstractWe present a new clustering algorithm for global illumination in complex environments. The new algorithm extends provious work on clustering for radiosity to allow for nondiffuse (glossy) reflectors. We represent clusters as points with directional distributions of outgoing and incoming radiance and importance, and we derive an error bound for transfers between these clusters. The algorithm groups input surfaces into a hierarchy of clusters, and then permits clusters to interact only if the error bound is below an acceptable tolerance. We show that the algorithm is asymptotically more efficient than previous clustering algorithms even when restricted to ideally diffuse environments. Finally, we demonstrate the performance of our method on two complex glossy environments. Per H. Christensen, Dani Lischinski, Eric J. Stollnitz, David Salesin |
ACM Trans. Graph. | 2 |
| 1996 | Fast Rendering of Complex Environments Using a Spatial Hierarchy
Bradford L. Chamberlain, Tony DeRose, Dani Lischinski, David Salesin |
Graphics Interface | 3 |
| 1996 | Scale-Dependent Reproduction of Pen-and-Ink IllustrationsabstractThis paper describes a representation for pen-and-ink illustrations that allows the creation of high-fidelity illustrations at any scale or resolution.We represent a pen-and-ink illustration as a low-resolution grey-scale image, augmented by a set of discontinuity segments, along with a stroke texture.To render an illustration at a particular scale, we first rescale the grey-scale image to the desired size and then hatch the resulting image with pen-and-ink strokes.The main technical contribution of the paper is a new reconstruction algorithm that magnifies the low-resolution image while keeping the resulting image sharp along discontinuities. Michael Salisbury, Corin R. Anderson, Dani Lischinski, David Salesin |
SIGGRAPH | 3 |
| 1996 | Hierarchical Image Caching for Accelerated Walkthroughs of Complex EnvironmentsabstractWe present a new method for accelerating walkthroughs of geometrically complex static scenes. As a preprocessing step, our method constructs a BSP-tree that hierarchically partitions the geometric primitives in the scene. In the course of a walkthrough, images of nodes at various levels of the hierarchy are cached for reuse in subsequent frames. A cached image is applied as a texture map to a single quadrilateral that is drawn instead of the geometry contained in the corresponding node. Visual artifacts are kept under control by using an error metric that quantifies the discrepancy between the appearance of the geometry contained in a node and the cached image. The new method is shown to achieve significant speedups for a walkthrough of a complex outdoor scene, with little or no loss in rendering quality. 1 Introduction Interactive visualization of extremely complex geometric environments is becoming an increasingly important application of computer graphics. Advances in the throughp... Jonathan Shade, Dani Lischinski, David Salesin, Tony DeRose, John M. Snyder |
SIGGRAPH | 2 |
| 1994 | Bounds and error estimates for radiosityabstractWe present a method for determining a posteriori bounds and estimates for local and total errors in radiosity solutions. The ability to obtain bounds and estimates for the total error is crucial for reliably judging the acceptability of a solution. Realistic estimates of the local error improve the efficiency of adaptive radiosity algorithms, such as hierarchical radiosity, by indicating where adaptive refinement is necessary. First, we describe a hierarchical radiosity algorithm that computes conservative lower and upper bounds on the exact radiosity function, as well as on the approximate solution. These bounds account for the propagation of errors due to interreflections, and provide a conservative upper bound on the error. We also describe a non-conservative version of the same algorithm that is capable of computing tighter bounds, from which more realistic error estimates can be obtained. Finally, we derive an expression for the effect of a particular interaction on the total erro... Dani Lischinski, Brian E. Smits, Donald P. Greenberg |
SIGGRAPH | 1 |
| 1993 | Combining hierarchical radiosity and discontinuity meshingabstractWe introduce a new approach for the computation of viewindependent solutions to the diffuse global illumination problem in polyhedral environments. The approach combines ideas from hierarchical radiosity and discontinuity meshing to yield solutions that are accurate both numerically and visually. First, we describe a modified hierarchical radiosity algorithm that uses a discontinuitydriven subdivision strategy to achieve better numerical accuracyand faster convergence. Second, we present a new algorithm based on discontinuity meshing that uses the hierarchical solution to reconstruct an object-space approximation to the radiance function that is visually accurate. Our results show significant improvements over both hierarchical radiosity and discontinuity meshing algorithms. CR Categories and Subject Descriptors: I.3.3---[Computer Graphics]: Picture/Image Generation; I.3.7---[Computer Graphics ]: Three-Dimensional Graphics and Realism. Additional Key Words and Phrases: diffuse refle... Dani Lischinski, Filippo Tampieri, Donald P. Greenberg |
SIGGRAPH | 1 |
| 1990 | Improved techniques for ray tracing parametric surfaces
Dani Lischinski, Jakob Gonczarowski |
Vis. Comput. | 1 |