VLDB 2026 Research / reviewers in the wild / expert
Lihi Zelnik-Manor
dblp:z/LihiZelnikManor · also Lihi Zelnik
· DBLP profile ↗
61ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0002-7930-8985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 12 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and ManipulationabstractAdvancements in text-to-image diffusion models have led to significant progress in fast 3D content creation. One common approach is to generate a set of multi-view images of an object, and then reconstruct it into a 3D model. However, this approach bypasses the use of a native 3D representation of the object and is hence prone to geometric artifacts and limited in controllability and manipulation capabilities. An alternative approach involves native 3D generative models that directly produce 3D representations. These models, however, are typically limited in their resolution, resulting in lower quality 3D objects. In this work, we bridge the quality gap between methods that directly generate 3D representations and ones that reconstruct 3D objects from multi-view images. We introduce a multi-view to multi-view diffusion model called Sharp-It, which takes a 3D consistent set of multi-view images rendered from a low-quality object and enriches its geometric details and texture. The diffusion model operates on the multi-view set in parallel, in the sense that it shares features across the generated views. A high-quality 3D model can then be reconstructed from the enriched multi-view set. By leveraging the advantages of both 2D and 3D approaches, our method offers an efficient and controllable method for high-quality 3D content creation. We demonstrate that Sharp-It enables various 3D applications, such as fast synthesis, editing, and controlled generation, while attaining high-quality assets. Yiftach Edelstein, Or Patashnik, Dana Cohen-Bar, Lihi Zelnik-Manor |
CVPR | 4 |
| 2024 | FreeAugment: Data Augmentation Search Across All Degrees of Freedom
Tom Bekor, Niv Nayman, Lihi Zelnik-Manor |
ECCV (27) | 3 |
| 2024 | Diverse Imagenet Models Transfer BetterabstractA commonly accepted hypothesis is that models with higher accuracy on Imagenet perform better on other downstream tasks, leading to much research dedicated to optimizing Imagenet accuracy. Recently this hypothesis has been challenged by evidence showing that self-supervised models transfer better than their supervised counterparts, despite their inferior Imagenet accuracy. This calls for identifying the additional factors, on top of Imagenet accuracy, that make models transferable. In this work we show that high diversity of the filters learnt by the model promotes transferability jointly with Imagenet accuracy. Encouraged by the recent transferability results of self-supervised models, we use a simple procedure to combine self-supervised and supervised pretraining and generate models with both high diversity and high accuracy, and as a result high transferability. We experiment with several architectures and multiple downstream tasks, including both single-label and multi-label classification. Niv Nayman, Avram Golbert, Asaf Noy, Lihi Zelnik-Manor |
WACV | 4 |
| 2023 | HUGO, a High-Resolution Tactile Emulator for Complex SurfacesabstractMany of our activities rely on tactile feedback perceived through mechanoreceptors in our skin. While visual and auditory devices provide immersive experiences, cutaneous feedback devices are typically limited in the range of sensations they provide and are hence usually used and tested on relatively simple synthetic surfaces. We present a device designed in a human-centered process, triggering the mechanoreceptors sensitive to pressure, low-frequency vibrations, and high-frequency vibrations, enabling one to experience touch of complex real-world surfaces. The device is based on a parallel manipulator and a pin-array, that operate simultaneously at 200Hz and emulate coarse and fine geometrical features, respectively. The decomposition into coarse and fine features, alongside the high operation frequency, enable simulation of virtual surfaces. This was corroborated via experiments on complex real-world surfaces via both a quantitative recognition test and a usability questionnaire. We believe that this design can be incorporated in numerous applications. Yair Herbst, Alon Wolf, Lihi Zelnik-Manor |
CHI | 3 |
| 2022 | BINAS: Bilinear Interpretable Neural Architecture Search
Niv Nayman, Yonathan Aflalo, Asaf Noy, Lihi Zelnik-Manor |
ACML | 4 |
| 2022 | Multi-label Classification with Partial Annotations using Class-aware Selective LossabstractLarge-scale multi-label classification datasets are commonly, and perhaps inevitably, partially annotated. That is, only a small subset of labels are annotated per sample. Different methods for handling the missing labels induce different properties on the model and impact its accuracy. In this work, we analyze the partial labeling problem, then propose a solution based on two key ideas. First, un-annotated labels should be treated selectively according to two probability quantities: the class distribution in the overall dataset and the specific label likelihood for a given data sample. We propose to estimate the class distribution using a dedicated temporary model, and we show its improved efficiency over a naive estimation computed using the dataset's partial annotations. Second, during the training of the target model, we emphasize the contribution of annotated labels over originally un-annotated labels by using a dedicated asymmetric loss. With our novel approach, we achieve state-of-the-art results on OpenImages dataset (e.g. reaching 87.3 mAP on V6). In addition, experiments conducted on LVIS and simulated-COCO demonstrate the effectiveness of our approach. Code is available at https://github.com/Alibaba-MIIL/PartialLabelingCSL. Emanuel Ben Baruch, Tal Ridnik, Itamar Friedman, Avi Ben-Cohen, Nadav Zamir, Asaf Noy, Lihi Zelnik-Manor |
CVPR | 7 |
| 2022 | PETA: Photo Albums Event Recognition using Transformers AttentionabstractIn recent years the amounts of personal photos captured increased significantly, giving rise to new challenges in high-level multi-image understanding. Event recognition in personal photo albums presents one challenging scenario where life events are recognized from a disordered collection of images, including both relevant and irrelevant images. Event recognition in images also presents the challenge of high-level image understanding, as opposed to low-level image object classification. In absence of methods to analyze multiple inputs, previous methods adopted temporal mechanisms, including various forms of recurrent neural networks. However, their effective temporal window is local. In addition, they are not a natural choice given the disordered characteristic of photo albums. We address this gap with a tailor-made solution, combining the power of CNNs for image representation and transformers for album representation to perform global reasoning on image collection, offering a practical and efficient solution for photo albums event recognition. Our solution reaches state-of-the-art results on three prominent benchmarks, achieving above 90% mAP on all datasets. We further explore the related image-importance task in event recognition, demonstrating how the learned attentions correlate with the human-annotated importance for this subjective task, thus opening the door for new applications.1 Tamar Glaser, Emanuel Ben Baruch, Gilad Sharir, Nadav Zamir, Asaf Noy, Lihi Zelnik-Manor |
ICPR | 6 |
| 2021 | Semantic Diversity Learning for Zero-Shot Multi-label ClassificationabstractTraining a neural network model for recognizing multiple labels associated with an image, including identifying unseen labels, is challenging, especially for images that portray numerous semantically diverse labels. As challenging as this task is, it is an essential task to tackle since it represents many real-world cases, such as image retrieval of natural images. We argue that using a single embedding vector to represent an image, as commonly practiced, is not sufficient to rank both relevant seen and unseen labels accurately. This study introduces an end-to-end model training for multi-label zero-shot learning that supports the semantic diversity of the images and labels. We propose to use an embedding matrix having principal embedding vectors trained using a tailored loss function. In addition, during training, we suggest up-weighting in the loss function image samples presenting higher semantic diversity to encourage the diversity of the embedding matrix. Extensive experiments show that our proposed method improves the zero-shot model’s quality in tag-based image retrieval achieving SoTA results on several common datasets (NUS-Wide, COCO, Open Images). Avi Ben-Cohen, Nadav Zamir, Emanuel Ben Baruch, Itamar Friedman, Lihi Zelnik-Manor |
ICCV | 5 |
| 2021 | Asymmetric Loss For Multi-Label ClassificationabstractIn a typical multi-label setting, a picture contains on average few positive labels, and many negative ones. This positive-negative imbalance dominates the optimization process, and can lead to under-emphasizing gradients from positive labels during training, resulting in poor accuracy. In this paper, we introduce a novel asymmetric loss ("ASL"), which operates differently on positive and negative samples. The loss enables to dynamically down-weights and hard-thresholds easy negative samples, while also discarding possibly mislabeled samples. We demonstrate how ASL can balance the probabilities of different samples, and how this balancing is translated to better mAP scores. With ASL, we reach state-of-the-art results on multiple popular multi-label datasets: MS-COCO, Pascal-VOC, NUS-WIDE and Open Images. We also demonstrate ASL applicability for other tasks, such as single-label classification and object detection. ASL is effective, easy to implement, and does not increase the training time or complexity. Implementation is available at: https://github.com/Alibaba-MIIL/ASL. Tal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, Lihi Zelnik-Manor |
ICCV | 7 |
| 2021 | HardCoRe-NAS: Hard Constrained diffeRentiable Neural Architecture SearchabstractRealistic use of neural networks often requires adhering to multiple constraints on latency, energy and memory among others. A popular approach to find fitting networks is through constrained Neural Architecture Search (NAS), however, previous methods enforce the constraint only softly. Therefore, the resulting networks do not exactly adhere to the resource constraint and their accuracy is harmed. In this work we resolve this by introducing Hard Constrained diffeRentiable NAS (HardCoRe-NAS), that is based on an accurate formulation of the expected resource requirement and a scalable search method that satisfies the hard constraint throughout the search. Our experiments show that HardCoRe-NAS generates state-of-the-art architectures, surpassing other NAS methods, while strictly satisfying the hard resource constraints without any tuning required. Niv Nayman, Yonathan Aflalo, Asaf Noy, Lihi Zelnik-Manor |
ICML | 4 |
| 2020 | ASAP: Architecture Search, Anneal and PruneabstractAutomatic methods for Neural ArchitectureSearch (NAS) have been shown to produce state-of-the-art network models, yet, their main drawback is the computational complexity of the search process. As some primal methods optimized over a discrete search space, thousands of days of GPU were required for convergence. A recent approach is based on constructing a differentiable search space that enables gradient-based optimization, thus reducing the search time to a few days. While successful, such methods still include some incontinuous steps, e.g., the pruning of many weak connections at once. In this paper, we propose a differentiable search space that allows the annealing of architecture weights, while gradually pruning inferior operations, thus the search converges to a single output network in a continuous manner. Experiments on several vision datasets demonstrate the effectiveness of our method with respect to the search cost, accuracy and the memory footprint of the achieved model. Asaf Noy, Niv Nayman, Tal Ridnik, Nadav Zamir, Sivan Doveh, Itamar Friedman, Raja Giryes, Lihi Zelnik-Manor |
AISTATS | 8 |
| 2020 | Graph Embedded Pose Clustering for Anomaly DetectionabstractWe propose a new method for anomaly detection of human actions. Our method works directly on human pose graphs that can be computed from an input video sequence. This makes the analysis independent of nuisance parameters such as viewpoint or illumination. We map these graphs to a latent space and cluster them. Each action is then represented by its soft-assignment to each of the clusters. This gives a kind of ”bag of words” representation to the data, where every action is represented by its similarity to a group of base action-words. Then, we use a Dirichlet process based mixture, that is useful for handling proportional data such as our soft-assignment vectors, to determine if an action is normal or not. We evaluate our method on two types of data sets. The first is a fine-grained anomaly detection data set (e.g. ShanghaiTech) where we wish to detect unusual variations of some action. The second is a coarse-grained anomaly detection data set (e.g., a Kinetics-based data set) where few actions are considered normal, and every other action should be considered abnormal. Extensive experiments on the benchmarks show that our method performs considerably better than other state of the art methods. Amir Markovitz, Gilad Sharir, Itamar Friedman, Lihi Zelnik-Manor, Shai Avidan |
CVPR | 4 |
| 2020 | Compact Network Training for Person ReIDabstractThe task of person re-identification (ReID) has attracted growing attention in recent years leading to improved performance, albeit with little focus on real-world applications. Most SotA methods are based on heavy pre-trained models, e.g. ResNet50 (~25M parameters), which makes them less practical and more tedious to explore architecture modifications. In this study, we focus on a small-sized randomly initialized model that enables us to easily introduce architecture and training modifications suitable for person ReID. The outcomes of our study are a compact network and a fitting training regime. We show the robustness of the network by outperforming the SotA on both Market1501 and DukeMTMC. Furthermore, we show the representation power of our ReID network via SotA results on a different task of multi-object tracking. Hussam Lawen, Avi Ben-Cohen, Matan Protter, Itamar Friedman, Lihi Zelnik-Manor |
ICMR | 5 |
| 2019 | Adversarial Feedback LoopabstractThanks to their remarkable generative capabilities, GANs have gained great popularity, and are used abundantly in state-of-the-art methods and applications. In a GAN based model, a discriminator is trained to learn the real data distribution. To date, it has been used only for training purposes, where it's utilized to train the generator to provide real-looking outputs. In this paper we propose a novel method that makes an explicit use of the discriminator in test-time, in a feedback manner in order to improve the generator results. To the best of our knowledge it is the first time a discriminator is involved in test-time. We claim that the discriminator holds significant information on the real data distribution, that could be useful for test-time as well, a potential that has not been explored before. The approach we propose does not alter the conventional training stage. At test-time, however, it transfers the output from the generator into the discriminator, and uses feedback modules (convolutional blocks) to translate the features of the discriminator layers into corrections to the features of the generator layers, which are used eventually to get a better generator result. Our method can contribute to both conditional and unconditional GANs. As demonstrated by our experiments, it can improve the results of state-of-the-art networks for super-resolution, and image generation. Firas Shama, Roey Mechrez, Alon Shoshan, Lihi Zelnik-Manor |
ICCV | 4 |
| 2019 | Dynamic-Net: Tuning the Objective Without Re-Training for Synthesis TasksabstractOne of the key ingredients for successful optimization of modern CNNs is identifying a suitable objective. To date, the objective is fixed a-priori at training time, and any variation to it requires re-training a new network. In this paper we present a first attempt at alleviating the need for re-training. Rather than fixing the network at training time, we train a ``Dynamic-Net'' that can be modified at inference time. Our approach considers an ``objective-space'' as the space of all linear combinations of two objectives, and the Dynamic-Net is emulating the traversing of this objective-space at test-time, without any further training. We show that this upgrades pre-trained networks by providing an out-of-learning extension, while maintaining the performance quality. The solution we propose is fast and allows a user to interactively modify the network, in real-time, in order to obtain the result he/she desires. We show the benefits of such an approach via several different applications. Alon Shoshan, Roey Mechrez, Lihi Zelnik-Manor |
ICCV | 3 |
| 2019 | XNAS: Neural Architecture Search with Expert AdviceabstractThis paper introduces a novel optimization method for differential neural architecture search, based on the theory of prediction with expert advice. Its optimization criterion is well fitted for an architecture-selection, i.e., it minimizes the regret incurred by a sub-optimal selection of operations. Unlike previous search relaxations, that require hard pruning of architectures, our method is designed to dynamically wipe out inferior architectures and enhance superior ones. It achieves an optimal worst-case regret bound and suggests the use of multiple learning-rates, based on the amount of information carried by the backward gradients. Experiments show that our algorithm achieves a strong performance over several image classification datasets. Specifically, it obtains an error rate of 1.6% for CIFAR-10, 23.9% for ImageNet under mobile settings, and achieves state-of-the-art results on three additional datasets. Niv Nayman, Asaf Noy, Tal Ridnik, Itamar Friedman, Rong Jin 0001, Lihi Zelnik-Manor |
NeurIPS | 6 |
| 2019 | Partial correspondence of 3D shapes using properties of the nearest-neighbor field
Nadav Yehonatan Arbel, Ayellet Tal, Lihi Zelnik-Manor |
Comput. Graph. | 3 |
| 2019 | Saliency driven image manipulation
Roey Mechrez, Eli Shechtman, Lihi Zelnik-Manor |
Mach. Vis. Appl. | 3 |
| 2018 | Maintaining Natural Image Statistics with the Contextual Loss
Roey Mechrez, Itamar Talmi, Firas Shama, Lihi Zelnik-Manor |
ACCV (3) | 4 |
| 2018 | Modifying Non-Local Variations Across Multiple ViewsabstractWe present an algorithm for modifying small non-local variations between repeating structures and patterns in multiple images of the same scene. The modification is consistent across views, even-though the images could have been photographed from different view points and under different lighting conditions. We show that when modifying each image independently the correspondence between them breaks and the geometric structure of the scene gets distorted. Our approach modifies the views while maintaining correspondence, hence, we succeed in modifying appearance and structure variations consistently. We demonstrate our methods on a number of challenging examples, photographed in different lighting, scales and view points. Tal Tlusty, Tomer Michaeli, Tali Dekel, Lihi Zelnik-Manor |
CVPR | 4 |
| 2018 | The Contextual Loss for Image Transformation with Non-aligned Data
Roey Mechrez, Itamar Talmi, Lihi Zelnik-Manor |
ECCV (14) | 3 |
| 2018 | Saliency Driven Image ManipulationabstractHave you ever taken a picture only to find out that an unimportant background object ended up being overly salient? Or one of those team sports photos where your favorite player blends with the rest? Wouldn't it be nice if you could tweak these pictures just a little bit so that the distractor would be attenuated and your favorite player will stand-out among her peers? Manipulating images in order to control the saliency of objects is the goal of this paper. We propose an approach that considers the internal color and saliency properties of the image. It changes the saliency map via an optimization framework that relies on patch-based manipulation using only patches from within the same image to maintain its appearance characteristics. Comparing our method to previous ones shows significant improvement, both in the achieved saliency manipulation and in the realistic appearance of the resulting images. Roey Mechrez, Eli Shechtman, Lihi Zelnik-Manor |
WACV | 3 |
| 2017 | Photorealistic Style Transfer with Screened Poisson Equation
Roey Mechrez, Eli Shechtman, Lihi Zelnik-Manor |
BMVC | 3 |
| 2017 | Template Matching with Deformable Diversity SimilarityabstractWe propose a novel measure for template matching named Deformable Diversity Similarity - based on the diversity of feature matches between a target image window and the template. We rely on both local appearance and geometric information that jointly lead to a powerful approach for matching. Our key contribution is a similarity measure, that is robust to complex deformations, significant background clutter, and occlusions. Empirical evaluation on the most up-to-date benchmark shows that our method outperforms the current state-of-the-art in its detection accuracy while improving computational complexity. Itamar Talmi, Roey Mechrez, Lihi Zelnik-Manor |
CVPR | 3 |
| 2017 | SIFTing Through ScalesabstractScale invariant feature detectors often find stable scales in only a few image pixels. Consequently, methods for feature matching typically choose one of two extreme options: matching a sparse set of scale invariant features, or dense matching using arbitrary scales. In this paper, we turn our attention to the overwhelming majority of pixels, those where stable scales are not found by standard techniques. We ask, is scale-selection necessary for these pixels, when dense, scale-invariant matching is required and if so, how can it be achieved? We make the following contributions: (i) We show that features computed over different scales, even in low-contrast areas, can be different and selecting a single scale, arbitrarily or otherwise, may lead to poor matches when the images have different scales. (ii) We show that representing each pixel as a set of SIFTs, extracted at multiple scales, allows for far better matches than single-scale descriptors, but at a computational price. Finally, (iii) we demonstrate that each such set may be accurately represented by a low-dimensional, linear subspace. A subspace-to-point mapping may further be used to produce a novel descriptor representation, the Scale-Less SIFT (SLS), as an alternative to single-scale descriptors. These claims are verified by quantitative and qualitative tests, demonstrating significant improvements over existing methods. A preliminary version of this work appeared in [1] . Tal Hassner, Shay Filosof, Viki Mayzels, Lihi Zelnik-Manor |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Approximate nearest neighbor fields in videoabstractWe introduce RIANN (Ring Intersection Approximate Nearest Neighbor search), an algorithm for matching patches of a video to a set of reference patches in real-time. For each query, RIANN finds potential matches by intersecting rings around key points in appearance space. Its search complexity is reversely correlated to the amount of temporal change, making it a good fit for videos, where typically most patches change slowly with time. Experiments show that RIANN is up to two orders of magnitude faster than previous ANN methods, and is the only solution that operates in real-time. We further demonstrate how RIANN can be used for real-time video processing and provide examples for a range of real-time video applications, including colorization, denoising, and several artistic effects. Nir Ben-Zrihem, Lihi Zelnik-Manor |
CVPR | 2 |
| 2015 | Hot or Not: Exploring Correlations between Appearance and TemperatureabstractIn this paper we explore interactions between the appearance of an outdoor scene and the ambient temperature. By studying statistical correlations between image sequences from outdoor cameras and temperature measurements we identify two interesting interactions. First, semantically meaningful regions such as foliage and reflective oriented surfaces are often highly indicative of the temperature. Second, small camera motions are correlated with the temperature in some scenes. We propose simple scene-specific temperature prediction algorithms which can be used to turn a camera into a crude temperature sensor. We find that for this task, simple features such as local pixel intensities outperform sophisticated, global features such as from a semantically-trained convolutional neural network. Daniel Glasner, Pascal Fua, Todd E. Zickler, Lihi Zelnik-Manor |
ICCV | 4 |
| 2014 | How to Evaluate Foreground MapsabstractThe output of many algorithms in computer-vision is either non-binary maps or binary maps (e.g., salient object detection and object segmentation). Several measures have been suggested to evaluate the accuracy of these foreground maps. In this paper, we show that the most commonly-used measures for evaluating both non-binary maps and binary maps do not always provide a reliable evaluation. This includes the Area-Under-the-Curve measure, the Average-Precision measure, the F-measure, and the evaluation measure of the PASCAL VOC segmentation challenge. We start by identifying three causes of inaccurate evaluation. We then propose a new measure that amends these flaws. An appealing property of our measure is being an intuitive generalization of the F-measure. Finally we propose four meta-measures to compare the adequacy of evaluation measures. We show via experiments that our novel measure is preferable. Ran Margolin, Lihi Zelnik-Manor, Ayellet Tal |
CVPR | 2 |
| 2014 | OTC: A Novel Local Descriptor for Scene Classification
Ran Margolin, Lihi Zelnik-Manor, Ayellet Tal |
ECCV (7) | 2 |
| 2013 | What Makes a Patch Distinct?abstractWhat makes an object salient? Most previous work assert that distinctness is the dominating factor. The difference between the various algorithms is in the way they compute distinctness. Some focus on the patterns, others on the colors, and several add high-level cues and priors. We propose a simple, yet powerful, algorithm that integrates these three factors. Our key contribution is a novel and fast approach to compute pattern distinctness. We rely on the inner statistics of the patches in the image for identifying unique patterns. We provide an extensive evaluation and show that our approach outperforms all state-of-the-art methods on the five most commonly-used datasets. Ran Margolin, Ayellet Tal, Lihi Zelnik-Manor |
CVPR | 3 |
| 2013 | Learning Video Saliency from Human Gaze Using Candidate SelectionabstractDuring recent years remarkable progress has been made in visual saliency modeling. Our interest is in video saliency. Since videos are fundamentally different from still images, they are viewed differently by human observers. For example, the time each video frame is observed is a fraction of a second, while a still image can be viewed leisurely. Therefore, video saliency estimation methods should differ substantially from image saliency methods. In this paper we propose a novel method for video saliency estimation, which is inspired by the way people watch videos. We explicitly model the continuity of the video by predicting the saliency map of a given frame, conditioned on the map from the previous frame. Furthermore, accuracy and computation speed are improved by restricting the salient locations to a carefully selected candidate set. We validate our method using two gaze-tracked video datasets and show we outperform the state-of-the-art. Dmitry Rudoy, Dan B. Goldman, Eli Shechtman, Lihi Zelnik-Manor |
CVPR | 4 |
| 2013 | SIFTpack: A Compact Representation for Efficient SIFT MatchingabstractComputing distances between large sets of SIFT descriptors is a basic step in numerous algorithms in computer vision. When the number of descriptors is large, as is often the case, computing these distances can be extremely time consuming. In this paper we propose the SIFT pack: a compact way of storing SIFT descriptors, which enables significantly faster calculations between sets of SIFTs than the current solutions. SIFT pack can be used to represent SIFTs densely extracted from a single image or sparsely from multiple different images. We show that the SIFT pack representation saves both storage space and run time, for both finding nearest neighbors and for computing all distances between all descriptors. The usefulness of SIFT pack is also demonstrated as an alternative implementation for K-means dictionaries of visual words. Alexandra Gilinsky, Lihi Zelnik-Manor |
ICCV | 2 |
| 2013 | Video inlays: a system for user-friendly matchmoveabstractDigital editing technology is highly popular as it enables to easily change photos and add to them artificial objects. Conversely, video editing is still challenging and mainly left to the professionals. Even basic video manipulations involve complicated software tools that are typically not adopted by the amateur user. In this paper we propose a system that allows an amateur user to performs a basic matchmove by adding an inlay to a video. Our system does not require any previous experience and relies on a simple user interaction. We allow adding 3D objects and volumetric textures to virtually any video. We demonstrate the method's applicability on a variety of videos downloaded from the web. Dmitry Rudoy, Lihi Zelnik-Manor |
VRST | 2 |
| 2013 | Saliency for image manipulation
Ran Margolin, Lihi Zelnik-Manor, Ayellet Tal |
Vis. Comput. | 2 |
| 2012 | Icon scanning: Towards next generation QR codesabstractUndoubtedly, a key feature in the popularity of smartmobile devices is the numerous applications one can install. Frequently, we learn about an application we desire by seeing it on a review site, someone else's device, or a magazine. A user-friendly way to obtain this particular application could be by taking a snapshot of its corresponding icon and being directed automatically to its download link. Such a solution exists today for QR codes, which can be thought of as icons with a binary pattern. In this paper we extend this to App-icons and propose a complete system for automatic icon-scanning: it first detects the icon in a snapshot and then recognizes it. Icon scanning is a highly challenging problem due to the large variety of icons (~500K in App-Store) and background wallpapers. In addition, our system should further deal with the challenges introduced by taking pictures of a screen. Nevertheless, the novel solution proposed in this paper provides high detection and recognition rates. We test our complete icon-scanning system on icon snapshots taken by independent users, and search them within the entire set of icons in App-Store. Our success rates are high and improve significantly on other methods. Itamar Friedman, Lihi Zelnik-Manor |
CVPR | 2 |
| 2012 | On SIFTs and their scalesabstractScale invariant feature detectors often find stable scales in only a few image pixels. Consequently, methods for feature matching typically choose one of two extreme options: matching a sparse set of scale invariant features, or dense matching using arbitrary scales. In this paper we turn our attention to the overwhelming majority of pixels, those where stable scales are not found by standard techniques. We ask, is scale-selection necessary for these pixels, when dense, scale-invariant matching is required and if so, how can it be achieved? We make the following contributions: (i) We show that features computed over different scales, even in low-contrast areas, can be different; selecting a single scale, arbitrarily or otherwise, may lead to poor matches when the images have different scales. (ii) We show that representing each pixel as a set of SIFTs, extracted at multiple scales, allows for far better matches than single-scale descriptors, but at a computational price. Finally, (iii) we demonstrate that each such set may be accurately represented by a low-dimensional, linear subspace. A subspace-to-point mapping may further be used to produce a novel descriptor representation, the Scale-Less SIFT (SLS), as an alternative to single-scale descriptors. These claims are verified by quantitative and qualitative tests, demonstrating significant improvements over existing methods. Tal Hassner, Viki Mayzels, Lihi Zelnik-Manor |
CVPR | 3 |
| 2012 | Viewpoint Selection for Human Actions
Dmitry Rudoy, Lihi Zelnik-Manor |
Int. J. Comput. Vis. | 2 |
| 2012 | Unsupervised Learning of Categorical Segments in Image CollectionsabstractWhich one comes first: segmentation or recognition? We propose a unified framework for carrying out the two simultaneously and without supervision. The framework combines a flexible probabilistic model, for representing the shape and appearance of each segment, with the popular “bag of visual words” model for recognition. If applied to a collection of images, our framework can simultaneously discover the segments of each image and the correspondence between such segments, without supervision. Such recurring segments may be thought of as the “parts” of corresponding objects that appear multiple times in the image collection. Thus, the model may be used for learning new categories, detecting/classifying objects, and segmenting images, without using expensive human annotation. Marco Andreetto, Lihi Zelnik-Manor, Pietro Perona |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Context-Aware Saliency DetectionabstractWe propose a new type of saliency—context-aware saliency—which aims at detecting the image regions that represent the scene. This definition differs from previous definitions whose goal is to either identify fixation points or detect the dominant object. In accordance with our saliency definition, we present a detection algorithm which is based on four principles observed in the psychological literature. The benefits of the proposed approach are evaluated in two applications where the context of the dominant objects is just as essential as the objects themselves. In image retargeting, we demonstrate that using our saliency prevents distortions in the important regions. In summarization, we show that our saliency helps to produce compact, appealing, and informative summaries. Stas Goferman, Lihi Zelnik-Manor, Ayellet Tal |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Approximate Nearest Subspace SearchabstractSubspaces offer convenient means of representing information in many pattern recognition, machine vision, and statistical learning applications. Contrary to the growing popularity of subspace representations, the problem of efficiently searching through large subspace databases has received little attention in the past. In this paper, we present a general solution to the problem of Approximate Nearest Subspace search. Our solution uniformly handles cases where the queries are points or subspaces, where query and database elements differ in dimensionality, and where the database contains subspaces of different dimensions. To this end, we present a simple mapping from subspaces to points, thus reducing the problem to the well-studied Approximate Nearest Neighbor problem on points. We provide theoretical proofs of correctness and error bounds of our construction and demonstrate its capabilities on synthetic and real data. Our experiments indicate that an approximate nearest subspace can be located significantly faster than the nearest subspace, with little loss of accuracy. Ronen Basri, Tal Hassner, Lihi Zelnik-Manor |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Posing to the Camera: Automatic Viewpoint Selection for Human Actions
Dmitry Rudoy, Lihi Zelnik-Manor |
ACCV (4) | 2 |
| 2010 | Context-aware saliency detectionabstractWe propose a new type of saliency - context-aware saliency - which aims at detecting the image regions that represent the scene. This definition differs from previous definitions whose goal is to either identify fixation points or detect the dominant object. In accordance with our saliency definition, we present a detection algorithm which is based on four principles observed in the psychological literature. The benefits of the proposed approach are evaluated in two applications where the context of the dominant objects is just as essential as the objects themselves. In image retargeting we demonstrate that using our saliency prevents distortions in the important regions. In summarization we show that our saliency helps to produce compact, appealing, and informative summaries. Stas Goferman, Lihi Zelnik-Manor, Ayellet Tal |
CVPR | 2 |
| 2010 | Large Scale Max-Margin Multi-Label Classification with Priors
Bharath Hariharan, Lihi Zelnik-Manor, S. V. N. Vishwanathan, Manik Varma |
ICML | 2 |
| 2010 | Puzzle-like CollageabstractAbstract Collages have been a common form of artistic expression since their first appearance in China around 200 BC. Recently, with the advance of digital cameras and digital image editing tools, collages have gained popularity also as a summarization tool. This paper proposes an approach for automating collage construction, which is based on assembling regions of interest of arbitrary shape in a puzzle‐like manner. We show that this approach produces collages that are informative, compact, and eye‐pleasing. This is obtained by following artistic principles and assembling the extracted cutouts such that their shapes complete each other. Stas Goferman, Ayellet Tal, Lihi Zelnik-Manor |
Comput. Graph. Forum | 3 |
| 2007 | Approximate Nearest Subspace Search with Applications to Pattern RecognitionabstractLinear and affine subspaces are commonly used to describe appearance of objects under different lighting, viewpoint, articulation, and identity. A natural problem arising from their use is - given a query image portion represented as a point in some high dimensional space - find a subspace near to the query. This paper presents an efficient solution to the approximate nearest subspace problem for both linear and affine subspaces. Our method is based on a simple reduction to the problem of nearest point search, and can thus employ tree based search or locality sensitive hashing to find a near subspace. Further speedup may be achieved by using random projections to lower the dimensionality of the problem. We provide theoretical proofs of correctness and error bounds of our construction and demonstrate its capabilities on synthetic and real data. Our experiments demonstrate that an approximate nearest subspace can be located significantly faster than the exact nearest subspace, while at the same time it can find better matches compared to a similar search on points, in the presence of variations due to viewpoint, lighting etc. Ronen Basri, Tal Hassner, Lihi Zelnik-Manor |
CVPR | 3 |
| 2007 | Non-Parametric Probabilistic Image SegmentationabstractWe propose a simple probabilistic generative model for image segmentation. Like other probabilistic algorithms (such as EM on a mixture of Gaussians) the proposed model is principled, provides both hard and probabilistic cluster assignments, as well as the ability to naturally incorporate prior knowledge. While previous probabilistic approaches are restricted to parametric models of clusters (e.g., Gaussians) we eliminate this limitation. The suggested approach does not make heavy assumptions on the shape of the clusters and can thus handle complex structures. Our experiments show that the suggested approach outperforms previous work on a variety of image segmentation tasks. Marco Andreetto, Lihi Zelnik-Manor, Pietro Perona |
ICCV | 2 |
| 2006 | On Single-Sequence and Multi-Sequence Factorizations
Lihi Zelnik-Manor, Michal Irani |
Int. J. Comput. Vis. | 1 |
| 2006 | Multi-body Factorization with Uncertainty: Revisiting Motion Consistency
Lihi Zelnik-Manor, Moshe Machline, Michal Irani |
Int. J. Comput. Vis. | 1 |
| 2006 | Statistical Analysis of Dynamic ActionsabstractReal-world action recognition applications require the development of systems which are fast, can handle a large variety of actions without a priori knowledge of the type of actions, need a minimal number of parameters, and necessitate as short as possible learning stage. In this paper, we suggest such an approach. We regard dynamic activities as long-term temporal objects, which are characterized by spatio-temporal features at multiple temporal scales. Based on this, we design a simple statistical distance measure between video sequences which captures the similarities in their behavioral content. This measure is nonparametric and can thus handle a wide range of complex dynamic actions. Having a behavior-based distance measure between sequences, we use it for a variety of tasks, including: video indexing, temporal segmentation, and action-based video clustering. These tasks are performed without prior knowledge of the types of actions, their models, or their temporal extents. Lihi Zelnik-Manor, Michal Irani |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Beyond Pairwise ClusteringabstractWe consider the problem of clustering in domains where the affinity relations are not dyadic (pairwise), but rather triadic, tetradic or higher. The problem is an instance of the hypergraph partitioning problem. We propose a two-step algorithm for solving this problem. In the first step we use a novel scheme to approximate the hypergraph using a weighted graph. In the second step a spectral partitioning algorithm is used to partition the vertices of this graph. The algorithm is capable of handling hyperedges of all orders including order two, thus incorporating information of all orders simultaneously. We present a theoretical analysis that relates our algorithm to an existing hypergraph partitioning algorithm and explain the reasons for its superior performance. We report the performance of our algorithm on a variety of computer vision problems and compare it to several existing hypergraph partitioning algorithms. Sameer Agarwal 0001, Jongwoo Lim, Lihi Zelnik-Manor, Pietro Perona, David J. Kriegman, Serge J. Belongie |
CVPR (2) | 3 |
| 2005 | Hybrid Models for Human Motion RecognitionabstractProbabilistic models have been previously shown to be efficient and effective for modeling and recognition of human motion. In particular we focus on methods which represent the human motion model as a triangulated graph. Previous approaches learned models based just on positions and velocities of the body parts while ignoring their appearance. Moreover, a heuristic approach was commonly used to obtain translation invariance. In this paper we suggest an improved approach for learning such models and using them for human motion recognition. The suggested approach combines multiple cues, i.e., positions, velocities and appearance into both the learning and detection phases. Furthermore, we introduce global variables in the model, which can represent global properties such as translation, scale or view-point. The model is learned in an unsupervised manner from unlabelled data. We show that the suggested hybrid probabilistic model (which combines global variables, like translation, with local variables, like relative positions and appearances of body parts), leads to: (i) faster convergence of learning phase, (it) robustness to occlusions, and, (Hi) higher recognition rate. Claudio Fanti, Lihi Zelnik-Manor, Pietro Perona |
CVPR (1) | 2 |
| 2005 | Squaring the Circles in PanoramasabstractPictures taken by a rotating camera cover the viewing sphere surrounding the center of rotation. Having a set of images registered and blended on the sphere what is left to be done, in order to obtain a flat panorama, is projecting the spherical image onto a picture plane. This step is unfortunately not obvious - the surface of the sphere may not be flattened onto a page without some form of distortion. The objective of this paper is discussing the difficulties and opportunities that are connected to the projection from viewing sphere to image plane. We first explore a number of alternatives to the commonly used linear perspective projection. These are 'global' projections and do not depend on image content. We then show that multiple projections may coexist successfully in the same mosaic: these projections are chosen locally and depend on what is present in the pictures. We show that such multi-view projections can produce more compelling results than the global projections Lihi Zelnik-Manor, Gabriele Peters, Pietro Perona |
ICCV | 1 |
| 2005 | Minimal-Cut Model CompositionabstractConstructing new, complex models is often done by reusing parts of existing models, typically by applying a sequence of segmentation, alignment and composition operations. Segmentation, either manual or automatic, is rarely adequate for this task, since it is applied to each model independently, leaving it to the user to trim the models and determine where to connect them. In this paper we propose a new composition tool. Our tool obtains as input two models, aligned either manually or automatically, and a small set of constraints indicating which portions of the two models should be preserved in the final output. It then automatically negotiates the best location to connect the models, trimming and stitching them as required to produce a seamless result. We offer a method based on the graph theoretic minimal cut as a means of implementing this new tool. We describe a system intended for both expert and novice users, allowing easy and flexible control over the composition result. In addition, we show our method to be well suited for a variety of model processing applications such as model repair, hole filling, and piecewise rigid deformations. Tal Hassner, Lihi Zelnik-Manor, George Leifman, Ronen Basri |
SMI | 2 |
| 2004 | Temporal Factorization vs. Spatial Factorization
Lihi Zelnik-Manor, Michal Irani |
ECCV (2) | 1 |
| 2004 | Self-Tuning Spectral ClusteringabstractWe study a number of open issues in spectral clustering: (i) Selecting the appropriate scale of analysis, (ii) Handling multi-scale data, (iii) Cluster- ing with irregular background clutter, and, (iv) Finding automatically the number of groups. We first propose that a ‘local’ scale should be used to compute the affinity between each pair of points. This local scaling leads to better clustering especially when the data includes multiple scales and when the clusters are placed within a cluttered background. We further suggest exploiting the structure of the eigenvectors to infer automatically the number of groups. This leads to a new algorithm in which the final randomly initialized k-means stage is eliminated. Lihi Zelnik-Manor, Pietro Perona |
NIPS | 1 |
| 2003 | Degeneracies, Dependencies and their Implications in Multi-body and Multi-Sequence FactorizationsabstractThe body of work on multi-body factorization separates between objects whose motions are independent. In this work we show that in many cases objects moving with different 3D motions will be captured as a single object using these approaches. We analyze what causes these degeneracies between objects and suggest an approach for overcoming some of them. We further show that in the case of multiple sequences linear dependencies can supply information for temporal synchronization of sequences and for spatial matching of points across sequences. Lihi Zelnik-Manor, Michal Irani |
CVPR (2) | 1 |
| 2002 | Multiview Constraints on HomographiesabstractThe image motion of a planar surface between two camera views is captured by a homography (a 2D projective transformation). The homography depends on the intrinsic and extrinsic camera parameters, as well as on the 3D plane parameters. While camera parameters vary across different views, the plane geometry remains the same. Based on this fact, we derive linear subspace constraints on the relative homographies of multiple (/spl ges/ 2) planes across multiple views. The paper has three main contributions: 1) We show that the collection of all relative homographies (homologies) of a pair of planes across multiple views, spans a 4-dimensional linear subspace. 2) We show how this constraint can be extended to the case of multiple planes across multiple views. 3) We show that, for some restricted cases of camera motion, linear subspace constraints apply also to the set of homographies of a single plane across multiple views. All the results derived are true for uncalibrated cameras. The possible utility of these multiview constraints for improving homography estimation and for detecting nonrigid motions are also discussed. Lihi Zelnik-Manor, Michal Irani |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Event-Based Analysis of VideoabstractDynamic events can be regarded as long-term temporal objects, which are characterized by spatio-temporal features at multiple temporal scales. Based on this, we design a simple statistical distance measure between video sequences (possibly of different lengths) based on their behavioral content. This measure is non-parametric and can thus handle a wide range of dynamic events. We use this measure for isolating and clustering events within long continuous video sequences. This is done without prior knowledge of the types of events, their models, or their temporal extent. An outcome of such a clustering process is a temporal segmentation of long video sequences into event-consistent sub-sequences, and their grouping into event-consistent clusters. Our event representation and associated distance measure can also be used for event-based indexing into long video sequences, even when only one short example-clip is available. However, when multiple example-clips of the same event are available (either as a result of the clustering process, or given manually), these can be used to refine the event representation, the associated distance measure, and accordingly the quality of the detection and clustering process. Lihi Zelnik-Manor, Michal Irani |
CVPR (2) | 1 |
| 2000 | Multi-Frame Estimation of Planar MotionabstractTraditional plane alignment techniques are typically performed between pairs of frames. We present a method for extending existing two-frame planar motion estimation techniques into a simultaneous multi-frame estimation, by exploiting multi-frame subspace constraints of planar surfaces. The paper has three main contributions: 1) we show that when the camera calibration does not change, the collection of all parametric image motions of a planar surface in the scene across multiple frames is embedded in a low dimensional linear subspace; 2) we show that the relative image motion of multiple planar surfaces across multiple frames is embedded in a yet lower dimensional linear subspace, even with varying camera calibration; and 3) we show how these multi-frame constraints can be incorporated into simultaneous multi-frame estimation of planar motion, without explicitly recovering any 3D information, or camera calibration. The resulting multi-frame estimation process is more constrained than the individual two-frame estimations, leading to more accurate alignment, even when applied to small image regions. Lihi Zelnik-Manor, Michal Irani |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Multi-Frame Alignment of PlanesabstractTraditional plane alignment techniques are typically performed between pairs of frames. In this paper we present a method for extending existing two-frame planar-motion estimation techniques into a simultaneous multi-frame estimation, by exploiting multi-frame geometric constraints of planar surfaces. The paper has three main contributions: (i) we show that when the camera calibration does not change, the collection of all parametric image motions of a planar surface in the scene across multiple frames is embedded in a low dimensional linear subspace; (ii) we show that the relative image motion of multiple planar surfaces across multiple frames is embedded in a yet lower dimensional linear subspace, even with varying camera calibration; and (iii) we show how these multi-frame constraints can be incorporated into simultaneous multi-frame estimation of planar motion, without explicitly recovering any 3D information, or camera calibration. The resulting multi-frame estimation process is more constrained than the individual two-frame estimations, leading to more accurate alignment, even when applied to small image regions. Lihi Zelnik-Manor, Michal Irani |
CVPR | 1 |
| 1999 | Multi-View Subspace Constraints on HomographiesabstractThe motion of a planar surface between two camera views induces a homography. The homography depends on the camera intrinsic and extrinsic parameters, as well as on the 3D plane parameters. While camera parameters vary across different views, the plane geometry remains the same. Based on this fact, the paper derives linear subspace constraints on the relative motion of multiple (/spl ges/2) planes across multiple views. The paper has three main contributions. It shows that the collection of all relative homographies of a pair of planes (homologies) across multiple views, spans a 4-dimensional linear subspace. It shows how this constraint can be extended to the case of multiple planes across multiple views. It suggests two potential application areas which can benefit from these constraints: the accuracy of homography estimation can be improved by enforcing the multi-view subspace constraints; and violations of these multi-view constraints can be used as a cue for moving object detection. All the results derived in this paper are true for uncalibrated cameras. Lihi Zelnik-Manor, Michal Irani |
ICCV | 1 |