VLDB 2026 Research / reviewers in the wild / expert
Shai Avidan
dblp:02/617
· DBLP profile ↗
118ranked-venue papers
23as first author
24since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 18 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 88 · 15 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-Aware Gaussian Splatting Decompositionabstract3D Gaussian Splatting (3D-GS) enables efficient novel view synthesis, but treats all frequencies uniformly, making it difficult to separate coarse structure from fine detail. Recent works have started to exploit frequency signals, but lack explicit frequency decomposition of the 3D representation itself. We propose a frequency-aware decomposition that organizes 3D Gaussians into groups corresponding to Laplacian-pyramid subbands of the input images. Each group is trained with spatial frequency regularization to confine it to its target frequency, while higher-frequency bands use signed residual colors to capture fine details that may be missed by lower-frequency reconstructions. A progressive coarse-to-fine training schedule stabilizes the decomposition. Our method achieves state-of-the-art reconstruction quality and rendering speed among all LOD-capable methods. In addition to improved interpretability, our method enables dynamic level-of-detail rendering, progressive streaming, foveated rendering, promptable 3D focus, and artistic filtering. Our code will be made publicly available. Yishai Lavi, Leo Segre, Shai Avidan |
3DV | 3 |
| 2026 | Decoupling Shape and Texture in SAM-2 via Controlled Texture Replacement
Inbal Cohen, Boaz Meivar, Peihan Tu, Shai Avidan, Gal Oren 0001 |
WACV | 4 |
| 2025 | Memories of Forgotten ConceptsabstractDiffusion models dominate the space of text-to-image generation, yet they may produce undesirable outputs, including explicit content or private data. To mitigate this, concept ablation techniques have been explored to limit the generation of certain concepts. In this paper, we reveal that the erased concept information persists in the model and that erased concept images can be generated using the right latent. Utilizing inversion methods, we show that there exist latent seeds capable of generating high quality images of erased concepts. Moreover, we show that these latents have likelihoods that overlap with those of images outside the erased concept. We extend this to demonstrate that for every image from the erased concept set, we can generate many seeds that generate the erased concept. Given the vast space of latents capable of generating ablated concept images, our results suggest that fully erasing concept information may be intractable, highlighting possible vulnerabilities in current concept ablation techniques. Matan Rusanovsky, Shimon Malnick, Amir Jevnisek, Ohad Fried, Shai Avidan |
CVPR | 5 |
| 2025 | CapeX: Category-Agnostic Pose Estimation from Textual Point ExplanationabstractConventional 2D pose estimation models are constrained by their design to specific object categories. This limits their applicability to predefined objects. To overcome these limitations, category-agnostic pose estimation (CAPE) emerged as a solution. CAPE aims to facilitate keypoint localization for diverse object categories using a unified model, which can generalize from minimal annotated support images.
Recent CAPE works have produced object poses based on arbitrary keypoint definitions annotated on a user-provided support image. Our work departs from conventional CAPE methods, which require a support image, by adopting a text-based approach instead of the support image.
Specifically, we use a pose-graph, where nodes represent keypoints that are described with text. This representation takes advantage of the abstraction of text descriptions and the structure imposed by the graph.
Our approach effectively breaks symmetry, preserves structure, and improves occlusion handling.
We validate our novel approach using the MP-100 benchmark, a comprehensive dataset covering over 100 categories and 18,000 images. MP-100 is structured so that the evaluation categories are unseen during training, making it especially suited for CAPE. Under a 1-shot setting, our solution achieves a notable performance boost of 1.26\%, establishing a new state-of-the-art for CAPE. Additionally, we enhance the dataset by providing text description annotations for both training and testing. We also include alternative text annotations specifically for testing the model's ability to generalize across different textual descriptions, further increasing its value for future research. Our code and dataset are publicly available at https://github.com/matanr/capex. Matan Rusanovsky, Or Hirschorn, Shai Avidan |
ICLR | 3 |
| 2025 | Lightning-Fast Image Inversion and Editing for Text-to-Image Diffusion ModelsabstractDiffusion inversion is the problem of taking an image and a text prompt that describes it and finding a noise latent that would generate the exact same image.
Most current deterministic inversion techniques operate by approximately solving an implicit equation and may converge slowly or yield poor reconstructed images. We formulate the problem by finding the roots of an implicit equation and devlop a method to solve it efficiently. Our solution is based on Newton-Raphson (NR), a well-known technique in numerical analysis. We show that a vanilla application of NR is computationally infeasible while naively transforming it to a computationally tractable alternative tends to converge to out-of-distribution solutions, resulting in poor reconstruction and editing. We therefore derive an efficient guided formulation that fastly converges and provides high-quality reconstructions and editing. We showcase our method on real image editing with three popular open-sourced diffusion models: Stable Diffusion, SDXL-Turbo, and Flux with different deterministic schedulers. Our solution, **Guided Newton-Raphson Inversion**, inverts an image within 0.4 sec (on an A100 GPU) for few-step models (SDXL-Turbo and Flux.1),
opening the door for interactive image editing. We further show improved results in image interpolation and generation of rare objects. Dvir Samuel, Barak Meiri, Haggai Maron, Yoad Tewel, Nir Darshan, Shai Avidan, Gal Chechik, Rami Ben-Ari |
ICLR | 6 |
| 2025 | Optimize the Unseen - Fast NeRF Cleanup with Free Space PriorabstractNeural Radiance Fields (NeRF) have advanced photorealistic novel view synthesis, but their reliance on photometric reconstruction introduces artifacts, commonly known as "floaters". These artifacts degrade novel view quality, particularly in unseen regions where NeRF optimization is unconstrained. We propose a fast, post-hoc NeRF cleanup method that eliminates such artifacts by enforcing a Free Space Prior, ensuring that unseen regions remain empty while preserving the structure of observed areas. Unlike existing approaches that rely on Maximum Likelihood (ML) estimation or complex, data-driven priors, our method adopts a Maximum-a-Posteriori (MAP) approach with a simple yet effective global prior. This enables our method to clean artifacts in both seen and unseen areas, significantly improving novel view quality even in challenging scene regions. Our approach generalizes across diverse NeRF architectures and datasets while requiring no additional memory beyond the original NeRF. Compared to state-of-the-art cleanup methods, our method is 2.5x faster in inference and completes cleanup training in under 30 seconds. Our code will be made publicly available. Leo Segre, Shai Avidan |
NeurIPS | 2 |
| 2025 | Sifting Through the Haystack - Efficiently Finding Rare Animal Behaviors in Large-Scale DatasetsabstractIn the study of animal behavior, researchers often record long continuous videos, accumulating into large-scale datasets. However, the behaviors of interest are often rare compared to routine behaviors. This incurs a heavy cost on manual annotation, forcing users to sift through many samples before finding their needles. We propose a pipeline to efficiently sample rare behaviors from large datasets, enabling the creation of training datasets for rare behavior classifiers. Our method only needs an unlabeled animal pose or acceleration dataset as input and makes no assumptions regarding the type, number, or characteristics of the rare behaviors. Our pipeline is based on a recent graph-based anomaly detection model for human behavior, which we apply to this new data domain. It leverages anomaly scores to automatically label normal samples while directing human annotation efforts toward anomalies. In research data, anomalies may come from many different sources (e.g., signal noise versus true rare instances). Hence, the entire labeling budget is focused on the abnormal classes, letting the user review and label samples according to their needs. We tested our approach on three datasets of freely-moving animals, acquired in the laboratory and the field. We found that graph-based models are particularly useful when studying motion-based behaviors in animals, yielding good results while using a small labeling budget. Our method consistently outperformed traditional random sampling, offering an average improvement of 70% in performance and creating datasets even when the behavior of interest was only 0.02% of the data. Even when the performance gain was minor (e.g., when the behavior is not rare), our method still reduced the annotation effort by half11Our code is available at https://github.com/shir3bar/SiftingTheHaystack.. Shir Bar, Or Hirschorn, Roi Holzman, Shai Avidan |
WACV | 4 |
| 2024 | Optimize & Reduce: A Top-Down Approach for Image VectorizationabstractVector image representation is a popular choice when editability and flexibility in resolution are desired. However, most images are only available in raster form, making raster-to-vector image conversion (vectorization) an important task. Classical methods for vectorization are either domain-specific or yield an abundance of shapes which limits editability and interpretability. Learning-based methods, that use differentiable rendering, have revolutionized vectorization, at the cost of poor generalization to out-of-training distribution domains, and optimization-based counterparts are either slow or produce non-editable and redundant shapes. In this work, we propose Optimize & Reduce (O&R), a top-down approach to vectorization that is both fast and domain-agnostic. O&R aims to attain a compact representation of input images by iteratively optimizing Bezier curve parameters and significantly reducing the number of shapes, using a devised importance measure. We contribute a benchmark of five datasets comprising images from a broad spectrum of image complexities - from emojis to natural-like images. Through extensive experiments on hundreds of images, we demonstrate that our method is domain agnostic and outperforms existing works in both reconstruction and perceptual quality for a fixed number of shapes. Moreover, we show that our algorithm is x10 faster than the state-of-the-art optimization-based method. Our code is publicly available: https://github.com/ajevnisek/optimize-and-reduce Or Hirschorn, Amir Jevnisek, Shai Avidan |
AAAI | 3 |
| 2024 | A Graph-Based Approach for Category-Agnostic Pose Estimation
Or Hirschorn, Shai Avidan |
ECCV (63) | 2 |
| 2024 | VF-NeRF: Viewshed Fields for Rigid NeRF Registration
Leo Segre, Shai Avidan |
ECCV (60) | 2 |
| 2024 | Taming Normalizing FlowsabstractWe propose an algorithm for taming Normalizing Flow models — changing the probability that the model will produce a specific image or image category. We focus on Normalizing Flows because they can calculate the exact generation probability likelihood for a given image. We demonstrate taming using models that generate human faces, a subdomain with many interesting privacy and bias considerations. Our method can be used in the context of privacy, e.g., removing a specific person from the output of a model, and also in the context of debiasing by forcing a model to output specific image categories according to a given distribution. Taming is achieved with a fast fine-tuning process without retraining from scratch, achieving the goal in a matter of minutes. We evaluate our method qualitatively and quantitatively, showing that the generation quality remains intact, while the desired changes are applied. Shimon Malnick, Shai Avidan, Ohad Fried |
WACV | 2 |
| 2024 | Hyperspectral image dynamic range reconstruction using deep neural network-based denoising methodsabstractAbstract Hyperspectral (HS) measurement is among the most useful tools in agriculture for early disease detection. However, the cost of HS cameras that can perform the desired detection tasks is prohibitive-typically fifty thousand to hundreds of thousands of dollars. In a previous study at the Agricultural Research Organization’s Volcani Institute (Israel), a low-cost, high-performing HS system was developed which included a point spectrometer and optical components. Its main disadvantage was long shooting time for each image. Shooting time strongly depends on the predetermined integration time of the point spectrometer. While essential for performing monitoring tasks in a reasonable time, shortening integration time from a typical value in the range of 200 ms to the 10 ms range results in deterioration of the dynamic range of the captured scene. In this work, we suggest correcting this by learning the transformation from data measured with short integration time to that measured with long integration time. Reduction of the dynamic range and consequent low SNR were successfully overcome using three developed deep neural networks models based on a denoising auto-encoder, DnCNN and LambdaNetworks architectures as a backbone. The best model was based on DnCNN using a combined loss function of $$\ell _{2}$$ ℓ 2 and Kullback–Leibler divergence on images with 20 consecutive channels. The full spectrum of the model achieved a mean PSNR of 30.61 and mean SSIM of 0.9, showing total improvement relatively to the 10 ms measurements’ mean PSNR and mean SSIM values by 60.43% and 94.51%, respectively. Loran Cheplanov, Shai Avidan, David J. Bonfil, Iftach Klapp |
Mach. Vis. Appl. | 2 |
| 2023 | SCOOP: Self-Supervised Correspondence and Optimization-Based Scene FlowabstractScene flow estimation is a long-standing problem in computer vision, where the goal is to find the 3D motion of a scene from its consecutive observations. Recently, there have been efforts to compute the scene flow from 3D point clouds. A common approach is to train a regression model that consumes source and target point clouds and outputs the per-point translation vector. An alternative is to learn point matches between the point clouds concurrently with regressing a refinement of the initial correspondence flow. In both cases, the learning task is very challenging since the flow regression is done in the free 3D space, and a typical solution is to resort to a large annotated synthetic dataset. We introduce SCOOP, a new method for scene flow estimation that can be learned on a small amount of data without employing ground-truth flow supervision. In contrast to previous work, we train a pure correspondence model focused on learning point feature representation and initialize the flow as the difference between a source point and its softly corresponding target point. Then, in the run-time phase, we directly optimize a flow refinement component with a self-supervised objective, which leads to a coherent and accurate flow field between the point clouds. Experiments on widespread datasets demonstrate the performance gains achieved by our method compared to existing leading techniques while using a fraction of the training data. Our code is publicly available11https://github.com/itailang/SCOOP . Itai Lang, Dror Aiger, Forrester Cole, Shai Avidan, Michael Rubinstein |
CVPR | 4 |
| 2023 | Normalizing Flows for Human Pose Anomaly DetectionabstractVideo anomaly detection is an ill-posed problem because it relies on many parameters such as appearance, pose, camera angle, background, and more. We distill the problem to anomaly detection of human pose, thus decreasing the risk of nuisance parameters such as appearance affecting the result. Focusing on pose alone also has the side benefit of reducing bias against distinct minority groups.Our model works directly on human pose graph sequences and is exceptionally lightweight (~1K parameters), capable of running on any machine able to run the pose estimation with negligible additional resources. We leverage the highly compact pose representation in a normalizing flows framework, which we extend to tackle the unique characteristics of spatio-temporal pose data and show its advantages in this use case.The algorithm is quite general and can handle training data of only normal examples as well as a supervised setting that consists of labeled normal and abnormal examples.We report state-of-the-art results on two anomaly detection benchmarks - the unsupervised ShanghaiTech dataset and the recent supervised UBnormal dataset. Code available at https://github.com/orhir/STG-NF. Or Hirschorn, Shai Avidan |
ICCV | 2 |
| 2023 | SAGA: Spectral Adversarial Geometric Attack on 3D MeshesabstractA triangular mesh is one of the most popular 3D data representations. As such, the deployment of deep neural networks for mesh processing is widely spread and is increasingly attracting more attention. However, neural networks are prone to adversarial attacks, where carefully crafted inputs impair the model’s functionality. The need to explore these vulnerabilities is a fundamental factor in the future development of 3D-based applications. Recently, mesh attacks were studied on the semantic level, where classifiers are misled to produce wrong predictions. Nevertheless, mesh surfaces possess complex geometric attributes beyond their semantic meaning, and their analysis often includes the need to encode and reconstruct the geometry of the shape.We propose a novel framework for a geometric adversarial attack on a 3D mesh autoencoder. In this setting, an adversarial input mesh deceives the autoencoder by forcing it to reconstruct a different geometric shape at its output. The malicious input is produced by perturbing a clean shape in the spectral domain. Our method leverages the spectral decomposition of the mesh along with additional mesh-related properties to obtain visually credible results that consider the delicacy of surface distortions1. Tomer Stolik, Itai Lang, Shai Avidan |
ICCV | 3 |
| 2022 | Learning ODIN
Amir Jevnisek, Shai Avidan |
BMVC | 2 |
| 2022 | A Two-Level Auto-Encoder for Distributed Stereo CodingabstractWe propose a new technique for stereo image compression that is based on Distributed Source Coding (DSC). In our setting, two cameras transmit their image back to a processing unit. Naively doing so requires each camera to compress and transmit its image independently. However, the images are correlated because they observe the same scene, and our goal is to take advantage of this fact. In our solution, one camera, assume the left camera, sends its image to the processing unit, as before. The right camera, on the other hand, transmits its image conditioned on the left image, even though the two cameras do not communicate. The processing unit can then decode the right image, using the left image. The solution is based on a two level Auto-Encoder (AE). During training, the first level AE learns a standard single image compression code. The second level AE further compresses the code of the right image, conditioned on the code of the left image. During inference, the left camera uses the first level AE to transmit its image to the processing unit. The right camera, on the other hand, uses the encoders of both levels to transmit its code to the processing unit. The processing unit uses the top level decoder to recover the left image, and the decoders of both levels, as well as the recovered left image, to recover the right image. The system achieves state of the art results in image compression on several popular datasets. Yuval Harel, Shai Avidan |
ICCP | 2 |
| 2022 | Aggregating Layers for Deepfake DetectionabstractThe increasing popularity of facial manipulation (Deepfakes) and synthetic face creation raises the need to develop robust forgery detection solutions. Crucially, most work in this domain assume that the Deepfakes in the test set come from the same Deepfake algorithms that were used for training the network. This is not how things work in practice. Instead, we consider the case where the network is trained on one Deepfake algorithm, and tested on Deepfakes generated by another algorithm. Typically, supervised techniques follow a pipeline of visual feature extraction from a deep backbone, followed by a binary classification head. Instead, our algorithm aggregates features extracted across all layers of one backbone network to detect a fake. We evaluate our approach on two domains of interest Deepfake detection and Synthetic image detection, and find that we achieve SOTA results. Amir Jevnisek, Shai Avidan |
ICPR | 2 |
| 2022 | Adversarial Mask: Real-World Universal Adversarial Attack on Face Recognition Models
Alon Zolfi, Shai Avidan, Yuval Elovici, Asaf Shabtai |
ECML/PKDD (3) | 2 |
| 2022 | Co-occurrence based texture synthesisabstractAs image generation techniques mature, there is a growing interest in explainable representations that are easy to understand and intuitive to manipulate. In this work, we turn to co-occurrence statistics, which have long been used for texture analysis, to learn a controllable texture synthesis model. We propose a fully convolutional generative adversarial network, conditioned locally on co-occurrence statistics, to generate arbitrarily large images while having local, interpretable control over texture appearance. To encourage fidelity to the input condition, we introduce a novel differentiable co-occurrence loss that is integrated seamlessly into our framework in an end-to-end fashion. We demonstrate that our solution offers a stable, intuitive, and interpretable latent representation for texture synthesis, which can be used to generate smooth texture morphs between different textures. We further show an interactive texture tool that allows a user to adjust local characteristics of the synthesized texture by directly using the co-occurrence values. Anna Darzi, Itai Lang, Ashutosh Taklikar, Hadar Averbuch-Elor, Shai Avidan |
Comput. Vis. Media | 5 |
| 2021 | DeepBBS: Deep Best Buddies for Point Cloud RegistrationabstractRecently, several deep learning approaches have been proposed for point cloud registration. These methods train a network to generate a representation that helps finding matching points in two 3D point clouds. Finding good matches allows them to calculate the transformation between the point clouds accurately. Two challenges of these techniques are dealing with occlusions and generalizing to objects of classes unseen during training. This work proposes DeepBBS, a novel method for learning a representation that takes into account the best buddy distance between points during training. Best Buddies (i.e., mutual nearest neighbors) are pairs of points nearest to each other. The Best Buddies criterion is a strong indication for correct matches that, in turn, leads to accurate registration. Our experiments show improved performance compared to previous methods. In particular, our learned representation leads to an accurate registration for partial shapes and in unseen categories. Our code is publicly available1 Itan Hezroni, Amnon Drory, Raja Giryes, Shai Avidan |
3DV | 4 |
| 2021 | DPC: Unsupervised Deep Point Correspondence via Cross and Self ConstructionabstractWe present a new method for real-time non-rigid dense correspondence between point clouds based on structured shape construction. Our method, termed Deep Point Correspondence (DPC), requires a fraction of the training data compared to previous techniques and presents better generalization capabilities. Until now, two main approaches have been suggested for the dense correspondence problem. The first is a spectral-based approach that obtains great results on synthetic datasets but requires mesh connectivity of the shapes and long inference processing time while being unstable in real-world scenarios. The second is a spatial approach that uses an encoder-decoder framework to regress an ordered point cloud for the matching alignment from an irregular input. Unfortunately, the decoder brings considerable disadvantages, as it requires a large amount of training data and struggles to generalize well in cross-dataset evaluations. DPC’s novelty lies in its lack of a decoder component. Instead, we use latent similarity and the input coordinates themselves to construct the point cloud and determine correspondence, replacing the coordinate regression done by the decoder. Extensive experiments show that our construction scheme leads to a performance boost in comparison to recent state-of-the-art correspondence methods. Our code is publicly available1. Itai Lang, Dvir Ginzburg, Shai Avidan, Dan Raviv |
3DV | 3 |
| 2021 | Geometric Adversarial Attacks and Defenses on 3D Point CloudsabstractDeep neural networks are prone to adversarial examples that maliciously alter the network’s outcome. Due to the increasing popularity of 3D sensors in safety-critical systems and the vast deployment of deep learning models for 3D point sets, there is a growing interest in adversarial attacks and defenses for such models. So far, the research has focused on the semantic level, namely, deep point cloud classifiers. However, point clouds are also widely used in a geometric-related form that includes encoding and reconstructing the geometry. In this work, we are the first to consider the problem of adversarial examples at a geometric level. In this setting, the question is how to craft a small change to a clean source point cloud that leads, after passing through an autoencoder model, to the reconstruction of a different target shape. Our attack is in sharp contrast to existing semantic attacks on 3D point clouds. While such works aim to modify the predicted label by a classifier, we alter the entire reconstructed geometry. Additionally, we demonstrate the robustness of our attack in the case of defense, where we show that remnant characteristics of the target shape are still present at the output after applying the defense to the adversarial input. Our code is publicly available1. Itai Lang, Uriel Kotlicki, Shai Avidan |
3DV | 3 |
| 2021 | Underwater Single Image Color Restoration Using Haze-Lines and a New Quantitative DatasetabstractUnderwater images suffer from color distortion and low contrast, because light is attenuated while it propagates through water. Attenuation under water varies with wavelength, unlike terrestrial images where attenuation is assumed to be spectrally uniform. The attenuation depends both on the water body and the 3D structure of the scene, making color restoration difficult. Unlike existing single underwater image enhancement techniques, our method takes into account multiple spectral profiles of different water types. By estimating just two additional global parameters: the attenuation ratios of the blue-red and blue-green color channels, the problem is reduced to single image dehazing, where all color channels have the same attenuation coefficients. Since the water type is unknown, we evaluate different parameters out of an existing library of water types. Each type leads to a different restored image and the best result is automatically chosen based on color distribution. We also contribute a dataset of 57 images taken in different locations. To obtain ground truth, we placed multiple color charts in the scenes and calculated its 3D structure using stereo imaging. This dataset enables a rigorous quantitative evaluation of restoration algorithms on natural images for the first time. Dana Berman, Deborah Levy, Shai Avidan, Tali Treibitz |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Proximity Preserving Binary Code Using Signed Graph-Cut
Inbal Lavi, Shai Avidan, Yoram Singer, Yacov Hel-Or |
AAAI | 2 |
| 2020 | Best Buddies Registration for Point Clouds
Amnon Drory, Tal Shomer, Shai Avidan, Raja Giryes |
ACCV (1) | 3 |
| 2020 | The Resistance to Label Noise in K-NN and DNN Depends on its Concentration
Amnon Drory, Oria Ratzon, Shai Avidan, Raja Giryes |
BMVC | 3 |
| 2020 | SampleNet: Differentiable Point Cloud SamplingabstractThere is a growing number of tasks that work directly on point clouds. As the size of the point cloud grows, so do the computational demands of these tasks. A possible solution is to sample the point cloud first. Classic sampling approaches, such as farthest point sampling (FPS), do not consider the downstream task. A recent work showed that learning a task-specific sampling can improve results significantly. However, the proposed technique did not deal with the non-differentiability of the sampling operation and offered a workaround instead. We introduce a novel differentiable relaxation for point cloud sampling that approximates sampled points as a mixture of points in the primary input cloud. Our approximation scheme leads to consistently good results on classification and geometry reconstruction applications. We also show that the proposed sampling method can be used as a front to a point cloud registration network. This is a challenging task since sampling must be consistent across two different point clouds for a shared downstream task. In all cases, our approach outperforms existing non-learned and learned sampling alternatives. Our code is publicly available. Itai Lang, Asaf Manor, Shai Avidan |
CVPR | 3 |
| 2020 | Graph Embedded Pose Clustering for Anomaly DetectionabstractWe propose a new method for anomaly detection of human actions. Our method works directly on human pose graphs that can be computed from an input video sequence. This makes the analysis independent of nuisance parameters such as viewpoint or illumination. We map these graphs to a latent space and cluster them. Each action is then represented by its soft-assignment to each of the clusters. This gives a kind of ”bag of words” representation to the data, where every action is represented by its similarity to a group of base action-words. Then, we use a Dirichlet process based mixture, that is useful for handling proportional data such as our soft-assignment vectors, to determine if an action is normal or not. We evaluate our method on two types of data sets. The first is a fine-grained anomaly detection data set (e.g. ShanghaiTech) where we wish to detect unusual variations of some action. The second is a coarse-grained anomaly detection data set (e.g., a Kinetics-based data set) where few actions are considered normal, and every other action should be considered abnormal. Extensive experiments on the benchmarks show that our method performs considerably better than other state of the art methods. Amir Markovitz, Gilad Sharir, Itamar Friedman, Lihi Zelnik-Manor, Shai Avidan |
CVPR | 5 |
| 2020 | Deep Image Compression Using Decoder Side Information
Sharon Ayzik, Shai Avidan |
ECCV (17) | 2 |
| 2020 | Unveiling Optical Properties in Underwater ImagesabstractThe appearance of underwater scenes is highly governed by the optical properties of the water (attenuation and scattering). However, most research effort in physics-based underwater image reconstruction methods is placed on devising image priors for estimating scene transmission, and less on estimating the optical properties. This limits the quality of the results. This work focuses on robust estimation of the water properties. First, as opposed to previous methods that used fixed values for attenuation, we estimate it from the color distribution in the image. Second, we estimate the veiling-light color from objects in the scene, contrary to looking at background pixels. We conduct an extensive qualitative and quantitative evaluation of our method vs. most recent methods on several datasets. As our estimation is more robust our method provides superior results including on challenging scenes. Yael Bekerman, Shai Avidan, Tali Treibitz |
ICCP | 2 |
| 2020 | NLDNet++: A Physics Based Single Image Dehazing NetworkabstractDeep learning methods for image dehazing achieve impressive results. Yet, the task of collecting ground truth hazy/dehazed image pairs to train the network is cumbersome. We propose to use Non-Local Image Dehazing (NLD), an existing physics based technique, to provide the dehazed image required to training a network. Upon close inspection, we find that NLD suffers from several shortcomings and propose novel extensions to improve it. The new method, termed NLD++, consists of 1) denoising the input image as pre-processing step to avoid noise amplification, 2) introducing a constrained optimization that respects physical constraints. NLD++ produces superior results to NLD at the expense of increased computational cost. To offset that, we propose NLDNet++, a fully convolutional network that is trained on pairs of hazy images and images dehazed by NLD++. This eliminates the need of existing deep learning methods that require hazy/dehazed image pairs that are difficult to obtain. We evaluate the performance of NLDNet++ on standard data sets and find it to compare favorably with existing methods. Iris Tal, Yael Bekerman, Avi Mor, Lior Knafo, Jonathan Alon, Shai Avidan |
ICCP | 6 |
| 2020 | Single Image Dehazing Using Haze-LinesabstractHaze often limits visibility and reduces contrast in outdoor images. The degradation varies spatially since it depends on the objects' distances from the camera. This dependency is expressed in the transmission coefficients, which control the attenuation. Restoring the scene radiance from a single image is a highly ill-posed problem, and thus requires using an image prior. Contrary to methods that use patch-based image priors, we propose an algorithm based on a non-local prior. The algorithm relies on the assumption that colors of a haze-free image are well approximated by a few hundred distinct colors, which form tight clusters in RGB space. Our key observation is that pixels in a given cluster are often non-local, i.e., spread over the entire image plane and located at different distances from the camera. In the presence of haze these varying distances translate to different transmission coefficients. Therefore, each color cluster in the clear image becomes a line in RGB space, that we term a haze-line. Using these haze-lines, our algorithm recovers the atmospheric light, the distance map and the haze-free image. The algorithm has linear complexity, requires no training, and performs well on a wide variety of images compared to other state-of-the-art methods. Dana Berman, Tali Treibitz, Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Learning to SampleabstractProcessing large point clouds is a challenging task. Therefore, the data is often sampled to a size that can be processed more easily. The question is how to sample the data? A popular sampling technique is Farthest Point Sampling (FPS). However, FPS is agnostic to a downstream application (classification, retrieval, etc.). The underlying assumption seems to be that minimizing the farthest point distance, as done by FPS, is a good proxy to other objective functions. We show that it is better to learn how to sample. To do that, we propose a deep network to simplify 3D point clouds. The network, termed S-NET, takes a point cloud and produces a smaller point cloud that is optimized for a particular task. The simplified point cloud is not guaranteed to be a subset of the original point cloud. Therefore, we match it to a subset of the original points in a post-processing step. We contrast our approach with FPS by experimenting on two standard data sets and show significantly better results for a variety of applications. Our code is publicly available. Oren Dovrat, Itai Lang, Shai Avidan |
CVPR | 3 |
| 2019 | Co-Occurrence Neural NetworkabstractConvolutional Neural Networks (CNNs) became a very popular tool for image analysis. Convolutions are fast to compute and easy to store, but they also have some limitations. First, they are shift-invariant and, as a result, they do not adapt to different regions of the image. Second, they have a fixed spatial layout, so small geometric deformations in the layout of a patch will completely change the filter response. For these reasons, we need multiple filters to handle the different parts and variations in the input. We augment the standard convolutional tools used in CNNs with a new filter that addresses both issues raised above. Our filter combines two terms, a spatial filter and a term that is based on the co-occurrence statistics of input values in the neighborhood. The proposed filter is differentiable and can therefore be packaged as a layer in CNN and trained using back-propagation. We show how to train the filter as part of the network and report results on several data sets. In particular, we replace a convolutional layer with hundreds of thousands of parameters with a Co-occurrence Layer consisting of only a few hundred parameters with minimal impact on accuracy. Irina Shevlev, Shai Avidan |
CVPR | 2 |
| 2018 | Matching Pixels Using Co-Occurrence StatisticsabstractWe propose a new error measure for matching pixels that is based on co-occurrence statistics. The measure relies on a co-occurrence matrix that counts the number of times pairs of pixel values co-occur within a window. The error incurred by matching a pair of pixels is inversely proportional to the probability that their values co-occur together, and not their color difference. This measure also works with features other than color, e.g. deep features. We show that this improves the state-of-the-art performance of template matching on standard benchmarks. We then propose an embedding scheme that maps the input image to an embedded image such that the Euclidean distance between pixel values in the embedded space resembles the co-occurrence statistics in the original space. This lets us run existing vision algorithms on the embedded images and enjoy the power of co-occurrence statistics for free. We demonstrate this on two algorithms, the Lucas-Kanade image registration and the Kernelized Correlation Filter (KCF) tracker. Experiments show that performance of each algorithm improves by about 10%. Rotal Kat, Roy Josef Jevnisek, Shai Avidan |
CVPR | 3 |
| 2018 | Guest Editorial: Vision and Computational Photography and Graphics
Radu Timofte, Luc Van Gool, Ming-Hsuan Yang 0001, Shai Avidan, Yasuyuki Matsushita, Qingxiong Yang |
Comput. Vis. Image Underst. | 4 |
| 2018 | Best-Buddies Similarity - Robust Template Matching Using Mutual Nearest NeighborsabstractWe propose a novel method for template matching in unconstrained environments. Its essence is the Best-Buddies Similarity (BBS), a useful, robust, and parameter-free similarity measure between two sets of points. BBS is based on counting the number of Best-Buddies Pairs (BBPs)-pairs of points in source and target sets that are mutual nearest neighbours, i.e., each point is the nearest neighbour of the other. BBS has several key features that make it robust against complex geometric deformations and high levels of outliers, such as those arising from background clutter and occlusions. We study these properties, provide a statistical analysis that justifies them, and demonstrate the consistent success of BBS on a challenging real-world dataset while using different types of features. Shaul Oron, Tali Dekel, Tianfan Xue, William T. Freeman, Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Color Restoration of Underwater Images
Dana Menaker, Tali Treibitz, Shai Avidan |
BMVC | 3 |
| 2017 | Co-occurrence FilterabstractCo-occurrence Filter (CoF) is a boundary preserving filter. It is based on the Bilateral Filter (BF) but instead of using a Gaussian on the range values to preserve edges it relies on a co-occurrence matrix. Pixel values that co-occur frequently in the image (i.e., inside textured regions) will have a high weight in the co-occurrence matrix. This, in turn, means that such pixel pairs will be averaged and hence smoothed, regardless of their intensity differences. On the other hand, pixel values that rarely co-occur (i.e., across texture boundaries) will have a low weight in the co-occurrence matrix. As a result, they will not be averaged and the boundary between them will be preserved. The CoF therefore extends the BF to deal with boundaries, not just edges. It learns co-occurrences directly from the image. We can achieve various filtering results by directing it to learn the co-occurrence matrix from a part of the image, or a different image. We give the definition of the filter, discuss how to use it with color images and show several use cases. Roy Josef Jevnisek, Shai Avidan |
CVPR | 2 |
| 2017 | Air-light estimation using haze-linesabstractOutdoor images taken in bad weather conditions, such as haze and fog, look faded and have reduced contrast. Recently there has been great success in single image dehazing, i.e., improving the visibility and restoring the colors from a single image. A crucial step in these methods is the calculation of the air-light color, the color of an area of the image with no objects in line-of-sight. We propose a new method for calculating the air-light. The method relies on the haze-lines prior that was recently introduced. This prior is based on the observation that the pixel values of a hazy image can be modeled as lines in RGB space that intersect at the air-light. We use Hough transform in RGB space to vote for the location of the air-light. We evaluate the proposed method on an existing dataset of real world images, as well as some synthetic and other real images. Our method performs on-par with current state-of-the-art techniques and is more computationally efficient. Dana Berman, Tali Treibitz, Shai Avidan |
ICCP | 3 |
| 2017 | Patch2Vec: Globally Consistent Image Patch RepresentationabstractAbstract Many image editing applications rely on the analysis of image patches. In this paper, we present a method to analyze patches by embedding them to a vector space, in which the Euclidean distance reflects patch similarity. Inspired by Word2Vec, we term our approachPatch2Vec. However, there is a significant difference between words and patches. Words have a fairly small and well defined dictionary. Image patches, on the other hand, have no such dictionary and the number of different patch types is not well defined. The problem is aggravated by the fact that each patch might contain several objects and textures. Moreover, Patch2Vec should be universal because it must be able to map never‐seen‐before texture to the vector space. The mapping is learned by analyzing the distribution of all natural patches. We use Convolutional Neural Networks (CNN) to learn Patch2Vec. In particular, we train a CNN on labeled images with a triplet‐loss objective function. The trained network encodes a given patch to a 128D vector. Patch2Vec is evaluated visually, qualitatively, and quantitatively. We then use several variants of an interactive single‐click image segmentation algorithm to demonstrate the power of our method. Ohad Fried, Shai Avidan, Daniel Cohen-Or |
Comput. Graph. Forum | 2 |
| 2017 | Detecting moving regions in CrowdCam images
Adi Dafni, Yael Moses, Shai Avidan, Tali Dekel |
Comput. Vis. Image Underst. | 3 |
| 2017 | Fast-Match: Fast Affine Template Matching
Simon Korman, Daniel Reichman 0001, Gilad Tsur, Shai Avidan |
Int. J. Comput. Vis. | 4 |
| 2016 | Non-local Image DehazingabstractHaze limits visibility and reduces image contrast in outdoor images. The degradation is different for every pixel and depends on the distance of the scene point from the camera. This dependency is expressed in the transmission coefficients, that control the scene attenuation and amount of haze in every pixel. Previous methods solve the single image dehazing problem using various patch-based priors. We, on the other hand, propose an algorithm based on a new, non-local prior. The algorithm relies on the assumption that colors of a haze-free image are well approximated by a few hundred distinct colors, that form tight clusters in RGB space. Our key observation is that pixels in a given cluster are often non-local, i.e., they are spread over the entire image plane and are located at different distances from the camera. In the presence of haze these varying distances translate to different transmission coefficients. Therefore, each color cluster in the clear image becomes a line in RGB space, that we term a haze-line. Using these haze-lines, our algorithm recovers both the distance map and the haze-free image. The algorithm is linear in the size of the image, deterministic and requires no training. It performs well on a wide variety of images and is competitive with other stateof-the-art methods. Dana Berman, Tali Treibitz, Shai Avidan |
CVPR | 3 |
| 2016 | Semi global boundary detection
Roy Josef Jevnisek, Shai Avidan |
Comput. Vis. Image Underst. | 2 |
| 2016 | Coherency Sensitive HashingabstractCoherency Sensitive Hashing (CSH) extends Locality Sensitivity Hashing (LSH) and PatchMatch to quickly find matching patches between two images. LSH relies on hashing, which maps similar patches to the same bin, in order to find matching patches. PatchMatch, on the other hand, relies on the observation that images are coherent, to propagate good matches to their neighbors in the image plane, using random patch assignment to seed the initial matching. CSH relies on hashing to seed the initial patch matching and on image coherence to propagate good matches. In addition, hashing lets it propagate information between patches with similar appearance (i.e., map to the same bin). This way, information is propagated much faster because it can use similarity in appearance space or neighborhood in the image plane. As a result, CSH is at least three to four times faster than PatchMatch and more accurate, especially in textured regions, where reconstruction artifacts are most noticeable to the human eye. We verified CSH on a new, large scale, data set of 133 image pairs and experimented on several extensions, including: k nearest neighbor search, the addition of rotation and matching three dimensional patches in videos. Simon Korman, Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Best-Buddies Similarity for robust template matchingabstractWe propose a novel method for template matching in unconstrained environments. Its essence is the Best-Buddies Similarity (BBS), a useful, robust, and parameter-free similarity measure between two sets of points. BBS is based on counting the number of Best-Buddies Pairs (BBPs)-pairs of points in source and target sets, where each point is the nearest neighbor of the other. BBS has several key features that make it robust against complex geometric deformations and high levels of outliers, such as those arising from background clutter and occlusions. We study these properties, provide a statistical analysis that justifies them, and demonstrate the consistent success of BBS on a challenging real-world dataset. Tali Dekel, Shaul Oron, Michael Rubinstein, Shai Avidan, William T. Freeman |
CVPR | 4 |
| 2015 | Inverting RANSAC: Global model detection via inlier rate estimationabstractThis work presents a novel approach for detecting inliers in a given set of correspondences (matches). It does so without explicitly identifying any consensus set, based on a method for inlier rate estimation (IRE). Given such an estimator for the inlier rate, we also present an algorithm that detects a globally optimal transformation. We provide a theoretical analysis of the IRE method using a stochastic generative model on the continuous spaces of matches and transformations. This model allows rigorous investigation of the limits of our IRE method for the case of 2D-translation, further giving bounds and insights for the more general case. Our theoretical analysis is validated empirically and is shown to hold in practice for the more general case of 2D-affinities. In addition, we show that the combined framework works on challenging cases of 2D-homography estimation, with very few and possibly noisy inliers, where RANSAC generally fails. Roee Litman, Simon Korman, Alexander M. Bronstein, Shai Avidan |
CVPR | 4 |
| 2015 | Peeking Template Matching for Depth ExtensionabstractWe propose a method that extends a given depth image into regions in 3D that are not visible from the point of view of the camera. The algorithm detects repeated 3D structures in the visible scene and suggests a set of 3D extension hypotheses, which are then combined together through a global 3D MRF discrete optimization. The recovered global 3D surface is consistent with both the input depth map and the hypotheses. A key component of this work is a novel 3D template matcher that is used to detect repeated 3D structure in the scene and to suggest the hypotheses. A unique property of this matcher is that it can handle depth uncertainty. This is crucial because the matcher is required to "peek around the corner", as it operates at the boundaries of the visible 3D scene where depth information is missing. The proposed matcher is fast and is guaranteed to find an approximation to the globally optimal solution. We demonstrate on real-world data that our algorithm is capable of completing a full 3D scene from a single depth image and can synthesize a full depth map from a novel viewpoint of the scene. In addition, we report results on an extensive synthetic set of 3D shapes, which allows us to evaluate the method both qualitatively and quantitatively. Simon Korman, Eyal Ofek, Shai Avidan |
ICCV | 3 |
| 2015 | Probably Approximately Symmetric: Fast Rigid Symmetry Detection With Global GuaranteesabstractAbstract We present a fast algorithm for global rigid symmetry detection with approximation guarantees. The algorithm is guaranteed to find the best approximate symmetry of a given shape, to within a user‐specified threshold, with very high probability. Our method uses a carefully designed sampling of the transformation space, where each transformation is efficiently evaluated using a sublinear algorithm. We prove that the density of the sampling depends on the total variation of the shape, allowing us to derive formal bounds on the algorithm's complexity and approximation quality. We further investigate different volumetric shape representations (in the form of truncated distance transforms), and in such a way control the total variation of the shape and hence the sampling density and the runtime of the algorithm. A comprehensive set of experiments assesses the proposed method, including an evaluation on the eight categories of the COSEG data set. This is the first large‐scale evaluation of any symmetry detection technique that we are aware of. Simon Korman, Roee Litman, Shai Avidan, Alexander M. Bronstein |
Comput. Graph. Forum | 3 |
| 2015 | Locally Orderless Tracking
Shaul Oron, Aharon Bar-Hillel, Dan Levi, Shai Avidan |
Int. J. Comput. Vis. | 4 |
| 2015 | Real-time tracking-with-detection for coping with viewpoint change
Shaul Oron, Aharon Bar-Hillel, Shai Avidan |
Mach. Vis. Appl. | 3 |
| 2014 | Extended Lucas-Kanade Tracking
Shaul Oron, Aharon Bar-Hillel, Shai Avidan |
ECCV (5) | 3 |
| 2014 | Photo Sequencing
Tali Dekel, Yael Moses, Shai Avidan |
Int. J. Comput. Vis. | 3 |
| 2013 | FasT-Match: Fast Affine Template MatchingabstractFast-Match is a fast algorithm for approximate template matching under 2D affine transformations that minimizes the Sum-of-Absolute-Differences (SAD) error measure. There is a huge number of transformations to consider but we prove that they can be sampled using a density that depends on the smoothness of the image. For each potential transformation, we approximate the SAD error using a sub linear algorithm that randomly examines only a small number of pixels. We further accelerate the algorithm using a branch-and-bound scheme. As images are known to be piecewise smooth, the result is a practical affine template matching algorithm with approximation guarantees, that takes a few seconds to run on a standard machine. We perform several experiments on three different datasets, and report very good results. To the best of our knowledge, this is the first template matching algorithm which is guaranteed to handle arbitrary 2D affine transformations. Simon Korman, Daniel Reichman 0001, Gilad Tsur, Shai Avidan |
CVPR | 4 |
| 2013 | Space-Time Tradeoffs in Photo SequencingabstractPhoto-sequencing is the problem of recovering the temporal order of a set of still images of a dynamic event, taken asynchronously by a set of uncalibrated cameras. Solving this problem is a first, crucial step for analyzing (or visualizing) the dynamic content of the scene captured by a large number of freely moving spectators. We propose a geometric based solution, followed by rank aggregation to the photo-sequencing problem. Our algorithm trades spatial certainty for temporal certainty. Whereas the previous solution proposed by [4] relies on two images taken from the same static camera to eliminate uncertainty in space, we drop the static-camera assumption and replace it with temporal information available from images taken from the same (moving) camera. Our method thus overcomes the limitation of the static-camera assumption, and scales much better with the duration of the event and the spread of cameras in space. We present successful results on challenging real data sets and large scale synthetic data (250 images). Tali Dekel, Yael Moses, Shai Avidan |
ICCV | 3 |
| 2013 | DCSH - Matching Patches in RGBD ImagesabstractWe extend patch based methods to work on patches in 3D space. We start with Coherency Sensitive Hashing (CSH), which is an algorithm for matching patches between two RGB images, and extend it to work with RGBD images. This is done by warping all 3D patches to a common virtual plane in which CSH is performed. To avoid noise due to warping of patches of various normals and depths, we estimate a group of dominant planes and compute CSH on each plane separately, before merging the matching patches. The result is DCSH - an algorithm that matches world (3D) patches in order to guide the search for image plane matches. An independent contribution is an extension of CSH, which we term Social-CSH. It allows a major speedup of the k nearest neighbor (kNN) version of CSH - its runtime growing linearly, rather than quadratic ally, in k. Social-CSH is used as a subcomponent of DCSH when many NNs are required, as in the case of image denoising. We show the benefits of using depth information to image reconstruction and image denoising, demonstrated on several RGBD images. Yaron Eshet, Simon Korman, Eyal Ofek, Shai Avidan |
ICCV | 4 |
| 2013 | Multiple histogram matchingabstractHistogram Matching (HM) is a common technique for finding a monotonic map between two histograms. However, HM cannot deal with cases where a single mapping is sought between two sets of histograms. This paper presents a novel technique that finds such a mapping in an optimal manner under various histograms distance measures. Dori Shapira, Shai Avidan, Yacov Hel-Or |
ICIP | 2 |
| 2013 | Stereo Seam Carving a Geometrically Consistent ApproachabstractImage retargeting algorithms attempt to adapt the image content to the screen without distorting the important objects in the scene. Existing methods address retargeting of a single image. In this paper, we propose a novel method for retargeting a pair of stereo images. Naively retargeting each image independently will distort the geometric structure and hence will impair the perception of the 3D structure of the scene. We show how to extend a single image seam carving to work on a pair of images. Our method minimizes the visual distortion in each of the images as well as the depth distortion. A key property of the proposed method is that it takes into account the visibility relations between pixels in the image pair (occluded and occluding pixels). As a result, our method guarantees, as we formally prove, that the retargeted pair is geometrically consistent with a feasible 3D scene, similar to the original one. Hence, the retargeted stereo pair can be viewed on a stereoscopic display or further processed by any computer vision algorithm. We demonstrate our method on a number of challenging indoor and outdoor stereo images. Tali Dekel, Yael Moses, Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Racing Bib Numbers Recognition
Idan Ben-ami, Tali Dekel, Shai Avidan |
BMVC | 3 |
| 2012 | Structure and motion from scene registrationabstractWe propose a method for estimating the 3D structure and the dense 3D motion (scene flow) of a dynamic nonrigid 3D scene, using a camera array. The core idea is to use a dense multi-camera array to construct a novel, dense 3D volumetric representation of the 3D space where each voxel holds an estimated intensity value and a confidence measure of this value. The problem of 3D structure and 3D motion estimation of a scene is thus reduced to a nonrigid registration of two volumes - hence the term ”Scene Registration”. Registering two dense 3D scalar volumes does not require recovering the 3D structure of the scene as a preprocessing step, nor does it require explicit reasoning about occlusions. From this nonrigid registration we accurately extract the 3D scene flow and the 3D structure of the scene, and successfully recover the sharp discontinuities in both time and space. We demonstrate the advantages of our method on a number of challenging synthetic and real data sets. Tali Dekel, Shai Avidan, Alexander Sorkine-Hornung, Wojciech Matusik |
CVPR | 2 |
| 2012 | Locally Orderless TrackingabstractLocally Orderless Tracking (LOT) is a visual tracking algorithm that automatically estimates the amount of local (dis)order in the object. This lets the tracker specialize in both rigid and deformable objects on-line and with no prior assumptions. We provide a probabilistic model of the object variations over time. The model is implemented using the Earth Mover's Distance (EMD) with two parameters that control the cost of moving pixels and changing their color. We adjust these costs on-line during tracking to account for the amount of local (dis)order in the object. We show LOT's tracking capabilities on challenging video sequences, both commonly used and new, demonstrating performance comparable to state-of-the-art methods. Shaul Oron, Aharon Bar-Hillel, Dan Levi, Shai Avidan |
CVPR | 4 |
| 2012 | Object retrieval and localization with spatially-constrained similarity measure and k-NN re-rankingabstractOne fundamental problem in object retrieval with the bag-of-visual words (BoW) model is its lack of spatial information. Although various approaches are proposed to incorporate spatial constraints into the BoW model, most of them are either too strict or too loose so that they are only effective in limited cases. We propose a new spatially-constrained similarity measure (SCSM) to handle object rotation, scaling, view point change and appearance deformation. The similarity measure can be efficiently calculated by a voting-based method using inverted files. Object retrieval and localization are then simultaneously achieved without post-processing. Furthermore, we introduce a novel and robust re-ranking method with the k-nearest neighbors of the query for automatically refining the initial search results. Extensive performance evaluations on six public datasets show that SCSM significantly outperforms other spatial models, while k-NN re-ranking outperforms most state-of-the-art approaches using query expansion. Xiaohui Shen, Zhe Lin 0001, Jonathan Brandt, Shai Avidan, Ying Wu 0001 |
CVPR | 4 |
| 2012 | Photo Sequencing
Tali Dekel, Yael Moses, Shai Avidan |
ECCV (6) | 3 |
| 2012 | TreeCANN - k-d Tree Coherence Approximate Nearest Neighbor Algorithm
Igor Olonetsky, Shai Avidan |
ECCV (4) | 2 |
| 2011 | Geometrically consistent stereo seam carvingabstractImage retargeting algorithms attempt to adapt the image content to the screen without distorting the important objects in the scene. Existing methods address retargeting of a single image. In this paper we propose a novel method for retargeting a pair of stereo images. Naively retargeting each image independently will distort the geometric structure and make it impossible to perceive the 3D structure of the scene. We show how to extend a single image seam carving to work on a pair of images. Our method minimizes the visual distortion in each of the images as well as the depth distortion. A key property of the proposed method is that it takes into account the visibility relations between pixels in the image pair (occluded and occluding pixels). As a result, our method guarantees, as we formally prove, that the retargeted pair is geometrically consistent with a feasible 3D scene, similar to the original one. Hence, the retargeted stereo pair can be viewed on a stereoscopic display or processed by any computer vision algorithm. We demonstrate our method on a number of challenging indoor and outdoor stereo images. Tali Dekel, Yael Moses, Shai Avidan |
ICCV | 3 |
| 2011 | Coherency Sensitive HashingabstractCoherency Sensitive Hashing (CSH) extends Locality Sensitivity Hashing (LSH) and PatchMatch to quickly find matching patches between two images. LSH relies on hashing, which maps similar patches to the same bin, in order to find matching patches. PatchMatch, on the other hand, relies on the observation that images are coherent, to propagate good matches to their neighbors, in the image plane. It uses random patch assignment to seed the initial matching. CSH relies on hashing to seed the initial patch matching and on image coherence to propagate good matches. In addition, hashing lets it propagate information between patches with similar appearance (i.e., map to the same bin). This way, information is propagated much faster because it can use similarity in appearance space or neighborhood in the image plane. As a result, CSH is at least three to four times faster than PatchMatch and more accurate, especially in textured regions, where reconstruction artifacts are most noticeable to the human eye. We verified CSH on a new, large scale, data set of 133 image pairs. Simon Korman, Shai Avidan |
ICCV | 2 |
| 2011 | Pause-and-play: automatically linking screencast video tutorials with applicationsabstractVideo tutorials provide a convenient means for novices to learn new software applications. Unfortunately, staying in sync with a video while trying to use the target application at the same time requires users to repeatedly switch from the application to the video to pause or scrub backwards to replay missed steps. We present Pause-and-Play, a system that helps users work along with existing video tutorials. Pause-and-Play detects important events in the video and links them with corresponding events in the target application as the user tries to replicate the depicted procedure. This linking allows our system to automatically pause and play the video to stay in sync with the user. Pause-and-Play also supports convenient video navigation controls that are accessible from within the target application and allow the user to easily replay portions of the video without switching focus out of the application. Finally, since our system uses computer vision to detect events in existing videos and leverages application scripting APIs to obtain real time usage traces, our approach is largely independent of the specific target application and does not require access or modifications to application source code. We have implemented Pause-and-Play for two target applications, Google SketchUp and Adobe Photoshop, and we report on a user study that shows our system improves the user experience of working with video tutorials. Suporn Pongnumkul, Mira Dontcheva, Wilmot Li, Jue Wang 0001, Lubomir D. Bourdev, Shai Avidan, Michael F. Cohen |
UIST | 6 |
| 2011 | CG2Real: Improving the Realism of Computer Generated Images Using a Large Collection of PhotographsabstractComputer-generated (CG) images have achieved high levels of realism. This realism, however, comes at the cost of long and expensive manual modeling, and often humans can still distinguish between CG and real images. We introduce a new data-driven approach for rendering realistic imagery that uses a large collection of photographs gathered from online repositories. Given a CG image, we retrieve a small number of real images with similar global structure. We identify corresponding regions between the CG and real images using a mean-shift cosegmentation algorithm. The user can then automatically transfer color, tone, and texture from matching regions to the CG image. Our system only uses image processing operations and does not require a 3D model of the scene, making it fast and easy to integrate into digital content creation workflows. Results of a user study show that our hybrid images appear more realistic than the originals. Micah K. Johnson, Kevin Dale, Shai Avidan, Hanspeter Pfister, William T. Freeman, Wojciech Matusik |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | A probabilistic image jigsaw puzzle solverabstractWe explore the problem of reconstructing an image from a bag of square, non-overlapping image patches, the jigsaw puzzle problem. Completing jigsaw puzzles is challenging and requires expertise even for humans, and is known to be NP-complete. We depart from previous methods that treat the problem as a constraint satisfaction problem and develop a graphical model to solve it. Each patch location is a node and each patch is a label at nodes in the graph. A graphical model requires a pairwise compatibility term, which measures an affinity between two neighboring patches, and a local evidence term, which we lack. This paper discusses ways to obtain these terms for the jigsaw puzzle problem. We evaluate several patch compatibility metrics, including the natural image statistics measure, and experimentally show that the dissimilarity-based compatibility - measuring the sum-of-squared color difference along the abutting boundary - gives the best results. We compare two forms of local evidence for the graphical model: a sparse-and-accurate evidence and a dense-and-noisy evidence. We show that the sparse-and-accurate evidence, fixing as few as 4 - 6 patches at their correct locations, is enough to reconstruct images consisting of over 400 patches. To the best of our knowledge, this is the largest puzzle solved in the literature. We also show that one can coarsely estimate the low resolution image from a bag of patches, suggesting that a bag of image patches encodes some geometric information about the original image. Taeg Sang Cho, Shai Avidan, William T. Freeman |
CVPR | 2 |
| 2010 | An eye for an eye: A single camera gaze-replacement methodabstractThe camera in video conference systems is typically positioned above, or below, the screen, causing the gaze of the users to appear misplaced. We propose an effective solution to this problem that is based on replacing the eyes of the user. This replacement, when done accurately, is enough to achieve a natural looking video. At an initialization stage the user is asked to look straight at the camera. We store these frames, then track the eyes accurately in the video sequence and replace the eyes, taking care of illumination and ghosting artifacts. We have tested the system on a large number of videos demonstrating the effectiveness of the proposed solution. Lior Wolf, Ziv Freund, Shai Avidan |
CVPR | 3 |
| 2010 | Multiple Hypothesis Video Segmentation from Superpixel Flows
Amelio Vázquez Reina, Shai Avidan, Hanspeter Pfister, Eric L. Miller 0001 |
ECCV (5) | 2 |
| 2010 | The Patch TransformabstractThe patch transform represents an image as a bag of overlapping patches sampled on a regular grid. This representation allows users to manipulate images in the patch domain, which then seeds the inverse patch transform to synthesize modified images. Possible modifications include the spatial locations of patches, the size of the output image, or the pool of patches from which an image is reconstructed. When no modifications are made, the inverse patch transform reduces to solving a jigsaw puzzle. The inverse patch transform is posed as a patch assignment problem on a Markov random field (MRF), where each patch should be used only once and neighboring patches should fit to form a plausible image. We find an approximate solution to the MRF using loopy belief propagation, introducing an approximation that encourages the solution to use each patch only once. The image reconstruction algorithm scales well with the total number of patches through label pruning. In addition, structural misalignment artifacts are suppressed through a patch jittering scheme that spatially jitters the assigned patches. We demonstrate the patch transform and its effectiveness on natural images. Taeg Sang Cho, Shai Avidan, William T. Freeman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Internet VisionabstractThe ten papers in this special issue focus on Internet vision. The goal is to provide the reader with a general sense of the research being conducted in Internet vision, defined as the intersection of computer vision and the Internet. This issue includes coverage of a number of significant advances in this field. Shai Avidan, Simon Baker, Ying Shan |
Proc. IEEE | 1 |
| 2010 | Infinite Images: Creating and Exploring a Large Photorealistic Virtual SpaceabstractWe present a system for generating “infinite” images from large collections of photos by means of transformed image retrieval. Given a query image, we first transform it to simulate how it would look if the camera moved sideways and then perform image retrieval based on the transformed image. We then blend the query and retrieved images to create a larger panorama. Repeating this process will produce an “infinite” image. The transformed image retrieval model is not limited to simple 2-D left/right image translation, however, and we show how to approximate other camera motions like rotation and forward motion/zoom-in using simple 2-D image transforms. We represent images in the database as a graph where each node is an image and different types of edges correspond to different types of geometric transformations simulating different camera motions. Generating infinite images is thus reduced to following paths in the image graph. Given this data structure we can also generate a panorama that connects two query images, simply by finding the shortest path between the two in the image graph. We call this option the “image taxi.” Our approach does not assume photographs are of a single real 3-D location, nor that they were taken at the same time. Instead, we organize the photos in themes, such as city streets or skylines and synthesize new virtual scenes by combining images from distinct but visually similar locations. There are a number of potential applications to this technology. It can be used to generate long panoramas as well as content aware transitions between reference images or video shots. Finally, the image graph allows users to interactively explore large photo collections for ideation, games, social interaction, and artistic purposes. Biliana Kaneva, Josef Sivic, Antonio Torralba 0001, Shai Avidan, William T. Freeman |
Proc. IEEE | 4 |
| 2009 | Mode-detection via median-shiftabstractMedian-shift is a mode seeking algorithm that relies on computing the median of local neighborhoods, instead of the mean. We further combine median-shift with Locality Sensitive Hashing (LSH) and show that the combined algorithm is suitable for clustering large scale, high dimensional data sets. In particular, we propose a new mode detection step that greatly accelerates performance. In the past, LSH was used in conjunction with mean shift only to accelerate nearest neighbor queries. Here we show that we can analyze the density of the LSH bins to quickly detect potential mode candidates and use only them to initialize the median-shift procedure. We use the median, instead of the mean (or its discrete counterpart - the medoid) because the median is more robust and because the median of a set is a point in the set. A median is well defined for scalars but there is no single agreed upon extension of the median to high dimensional data. We adopt a particular extension, known as the Tukey median, and show that it can be computed efficiently using random projections of the high dimensional data onto 1D lines, just like LSH, leading to a tightly integrated and efficient algorithm. Lior Shapira, Shai Avidan, Ariel Shamir |
ICCV | 2 |
| 2009 | Multi-operator media retargetingabstractContent aware resizing gained popularity lately and users can now choose from a battery of methods to retarget their media. However, no single retargeting operator performs well on all images and all target sizes. In a user study we conducted, we found that users prefer to combine seam carving with cropping and scaling to produce results they are satisfied with. This inspires us to propose an algorithm that combines different operators in an optimal manner. We define a resizing space as a conceptual multi-dimensional space combining several resizing operators, and show how a path in this space defines a sequence of operations to retarget media. We define a new image similarity measure, which we term Bi-Directional Warping (BDW), and use it with a dynamic programming algorithm to find an optimal path in the resizing space. In addition, we show a simple and intuitive user interface allowing users to explore the resizing space of various image sizes interactively. Using key-frames and interpolation we also extend our technique to retarget video, providing the flexibility to use the best combination of operators at different times in the sequence. Michael Rubinstein, Ariel Shamir, Shai Avidan |
ACM Trans. Graph. | 3 |
| 2008 | The patch transform and its applications to image editingabstractWe introduce the patch transform, where an image is broken into non-overlapping patches, and modifications or constraints are applied in the “patch domain”. A modified image is then reconstructed from the patches, subject to those constraints. When no constraints are given, the reconstruction problem reduces to solving a jigsaw puzzle. Constraints the user may specify include the spatial locations of patches, the size of the output image, or the pool of patches from which an image is reconstructed. We define terms in a Markov network to specify a good image reconstruction from patches: neighboring patches must fit to form a plausible image, and each patch should be used only once. We find an approximate solution to the Markov network using loopy belief propagation, introducing an approximation to handle the combinatorially difficult patch exclusion constraint. The resulting image reconstructions show the original image, modified to respect the user’s changes. We apply the patch transform to various image editing tasks and show that the algorithm performs well on real world images. Taeg Sang Cho, Moshe Butman, Shai Avidan, William T. Freeman |
CVPR | 3 |
| 2008 | Boundary snapping for robust image cutoutsabstractBoundary snapping is an interactive image cutout algorithm that requires a small number of user supplied control points, or landmarks, to infer the cutout contour. The key idea is to match the appearance of all points along the desired contour to the landmark points, where appearance is given by an intensity profile perpendicular to the boundary. An optimization process attempts to find a contour that maximizes the similarity score of its points with the landmarks. This approach works well in the typical case where the foreground and background differ in appearance, as well as in challenging cases where the subject is clearly perceived, but the regions on both sides of the boundary are similar and cannot be easily discriminated. By enabling the user to define the boundary points directly, the technique is not limited to boundaries that necessarily have to be the most salient or high gradient feature in the region. It can also be used for margin cutout around the boundary. The use of multiple control points along the boundary can handle spatially varying attributes as both foreground and background may change in appearance along the boundary. The final result is accurate, because it allows the user to enforce hard constraints on the boundary directly, at the expense of moderate user labor in positioning the landmark points. Finally, the algorithm is fast, works on a variety of images, and handles situations where the boundary is not obvious. Eyal Zadicario, Shai Avidan, Alon Shmueli, Daniel Cohen-Or |
CVPR | 2 |
| 2008 | Privacy Preserving Pattern ClassificationabstractWe give efficient and practical protocols for privacy preserving pattern classification that allow a client to have his data classified by a server, without revealing information to either party, other than the classification result. We illustrate the advantages of such a framework on several real-world scenarios and give secure protocols for several classifiers. Shai Avidan, Ariel Elbaz, Tal Malkin |
ICIP | 1 |
| 2008 | Light mixture estimation for spatially varying white balanceabstractWhite balance is a crucial step in the photographic pipeline. It ensures the proper rendition of images by eliminating color casts due to differing illuminants. Digital cameras and editing programs provide white balance tools that assume a single type of light per image, such as daylight. However, many photos are taken under mixed lighting. We propose a white balance technique for scenes with two light types that are specified by the user. This covers many typical situations involving indoor/outdoor or flash/ambient light mixtures. Since we work from a single image, the problem is highly underconstrained. Our method recovers a set of dominant material colors which allows us to estimate the local intensity mixture of the two light types. Using this mixture, we can neutralize the light colors and render visually pleasing images. Our method can also be used to achieve post-exposure relighting effects. Eugene Hsu, Tom Mertens, Sylvain Paris, Shai Avidan, Frédo Durand |
ACM Trans. Graph. | 4 |
| 2008 | Improved seam carving for video retargetingabstractVideo, like images, should support content aware resizing. We present video retargeting using an improved seam carving operator. Instead of removing 1D seams from 2D images we remove 2D seam manifolds from 3D space-time volumes. To achieve this we replace the dynamic programming method of seam carving with graph cuts that are suitable for 3D volumes. In the new formulation, a seam is given by a minimal cut in the graph and we show how to construct a graph such that the resulting cut is a valid seam. That is, the cut is monotonic and connected. In addition, we present a novel energy criterion that improves the visual quality of the retargeted images and videos. The original seam carving operator is focused on removing seams with the least amount of energy, ignoring energy that is introduced into the images and video by applying the operator. To counter this, the new criterion is looking forward in time - removing seams that introduce the least amount of energy into the retargeted result. We show how to encode the improved criterion into graph cuts (for images and video) as well as dynamic programming (for images). We apply our technique to images and videos and present results of various applications. Michael Rubinstein, Ariel Shamir, Shai Avidan |
ACM Trans. Graph. | 3 |
| 2007 | Statistics of Infrared ImagesabstractThe proliferation of low-cost infrared cameras gives us a new angle for attacking many unsolved vision problems by leveraging a larger range of the electromagnetic spectrum. A first step to utilizing these images is to explore the statistics of infrared images and compare them to the corresponding statistics in the visible spectrum. In this paper, we analyze the power spectra as well as the marginal and joint wavelet coefficient distributions of datasets of indoor and outdoor images. We note that infrared images have noticeably less texture indoors where temperatures are more homogenous. The joint wavelet statistics also show strong correlation between object boundaries in IR and visible images, leading to high potential for vision applications using a combined statistical model. Nigel J. W. Morris, Shai Avidan, Wojciech Matusik, Hanspeter Pfister |
CVPR | 2 |
| 2007 | Synthetic Aperture Tracking: Tracking through OcclusionsabstractOcclusion is a significant challenge for many tracking algorithms. Most current methods can track through transient occlusion, but cannot handle significant extended occlusion when the object's trajectory may change significantly. We present a method to track a 3D object through significant occlusion using multiple nearby cameras (e.g., a camera array). When an occluder and object are at different depths, different parts of the object are visible or occluded in each view due to parallax. By aggregating across these views, the method can track even when any individual camera observes very little of the target object. Implementation- wise, the methods are straightforward and build upon established single-camera algorithms. They do not require explicit modeling or reconstruction of the scene and enable tracking in complex, dynamic scenes with moving cameras. Analysis of accuracy and robustness shows that these methods are successful when upwards of '70% of the object is occluded in every camera view. To the best of our knowledge, this system is the first capable of tracking in the presence of such significant occlusion. Neel Joshi, Shai Avidan, Wojciech Matusik, David J. Kriegman |
ICCV | 2 |
| 2007 | Fast Pixel/Part Selection with Sparse EigenvectorsabstractWe extend the "Sparse LDA" algorithm of [7] with new sparsity bounds on 2-class separability and efficient partitioned matrix inverse techniques leading to 1000-fold speed-ups. This mitigates the 0(n4) scaling that has limited this algorithm's applicability to vision problems and also prioritizes the less-myopic backward elimination stage by making it faster than forward selection. Experiments include "sparse eigenfaces" and gender classification on FERET data as well as pixel/part selection for OCR on MNIST data using Bayesian (GP) classification. Sparse- LDA is an attractive alternative to the more demanding Automatic Relevance Determination. State-of-the-art recognition is obtained while discarding the majority of pixels in all experiments. Our sparse models also show a better fit to data in terms of the "evidence" or marginal likelihood. Bernard Moghaddam, Yair Weiss, Shai Avidan |
ICCV | 3 |
| 2007 | Centralized and Distributed Multi-view Correspondence
Shai Avidan, Yael Moses, Yoram Moses |
Int. J. Comput. Vis. | 1 |
| 2007 | Ensemble TrackingabstractWe consider tracking as a binary classification problem, where an ensemble of weak classifiers is trained online to distinguish between the object and the background. The ensemble of weak classifiers is combined into a strong classifier using AdaBoost. The strong classifier is then used to label pixels in the next frame as either belonging to the object or the background, giving a confidence map. The peak of the map and, hence, the new position of the object, is found using mean shift. Temporal coherence is maintained by updating the ensemble with new weak classifiers that are trained online during tracking. We show a realization of this method and demonstrate it on several video sequences. Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Seam carving for content-aware image resizingabstractEffective resizing of images should not only use geometric constraints, but consider the image content as well. We present a simple image operator called seam carving that supports content-aware image resizing for both reduction and expansion. A seam is an optimal 8-connected path of pixels on a single image from top to bottom, or left to right, where optimality is defined by an image energy function. By repeatedly carving out or inserting seams in one direction we can change the aspect ratio of an image. By applying these operators in both directions we can retarget the image to a new size. The selection and order of seams protect the content of the image, as defined by the energy function. Seam carving can also be used for image content enhancement and object removal. We support various visual saliency measures for defining the energy of an image, and can also include user input to guide the process. By storing the order of seams in an image we create multi-size images, that are able to continuously change in real time to fit a given size. Shai Avidan, Ariel Shamir |
ACM Trans. Graph. | 1 |
| 2006 | Fast Human Detection Using a Cascade of Histograms of Oriented GradientsabstractWe integrate the cascade-of-rejectors approach with the Histograms of Oriented Gradients (HoG) features to achieve a fast and accurate human detection system. The features used in our system are HoGs of variable-size blocks that capture salient features of humans automatically. Using AdaBoost for feature selection, we identify the appropriate set of blocks, from a large set of possible blocks. In our system, we use the integral image representation and a rejection cascade which significantly speed up the computation. For a 320 × 280 image, the system can process 5 to 30 frames per second depending on the density in which we scan the image, while maintaining an accuracy level similar to existing methods. Qiang Zhu 0006, Mei-Chen Yeh, Kwang-Ting Cheng, Shai Avidan |
CVPR (2) | 4 |
| 2006 | SpatialBoost: Adding Spatial Reasoning to AdaBoost
Shai Avidan |
ECCV (4) | 1 |
| 2006 | Blind Vision
Shai Avidan, Moshe Butman |
ECCV (3) | 1 |
| 2006 | Generalized spectral bounds for sparse LDAabstractWe present a discrete spectral framework for the sparse or cardinality-constrained solution of a generalized Rayleigh quotient. This NP-hard combinatorial optimization problem is central to supervised learning tasks such as sparse LDA, feature selection and relevance ranking for classification. We derive a new generalized form of the Inclusion Principle for variational eigenvalue bounds, leading to exact and optimal sparse linear discriminants using branch-and-bound search. An efficient greedy (approximate) technique is also presented. The generalization performance of our sparse LDA algorithms is demonstrated with real-world UCI ML benchmarks and compared to a leading SVM-based gene selection algorithm for cancer classification. Baback Moghaddam, Yair Weiss, Shai Avidan |
ICML | 3 |
| 2006 | Efficient Methods for Privacy Preserving Face DetectionabstractBob offers a face-detection web service where clients can submit their images for analysis. Alice would very much like to use the service, but is reluctant to reveal the content of her images to Bob. Bob, for his part, is reluctant to release his face detector, as he spent a lot of time, energy and money constructing it. Secure Multi- Party computations use cryptographic tools to solve this problem without leaking any information. Unfortunately, these methods are slow to compute and we intro- duce a couple of machine learning techniques that allow the parties to solve the problem while leaking a controlled amount of information. The first method is an information-bottleneck variant of AdaBoost that lets Bob find a subset of features that are enough for classifying an image patch, but not enough to actually recon- struct it. The second machine learning technique is active learning that allows Alice to construct an online classifier, based on a small number of calls to Bob’s face detector. She can then use her online classifier as a fast rejector before using a cryptographically secure classifier on the remaining image patches. Shai Avidan, Moshe Butman |
NIPS | 1 |
| 2006 | Natural video matting using camera arraysabstractWe present an algorithm and a system for high-quality natural video matting using a camera array. The system uses high frequencies present in natural scenes to compute mattes by creating a synthetic aperture image that is focused on the foreground object, which reduces the variance of pixels reprojected from the foreground while increasing the variance of pixels reprojected from the background. We modify the standard matting equation to work directly with variance measurements and show how these statistics can be used to construct a trimap that is later upgraded to an alpha matte. The entire process is completely automatic, including an automatic method for focusing the synthetic aperture image on the foreground object and an automatic method to compute the trimap and the alpha matte. The proposed algorithm is very efficient and has a per-pixel running time that is linear in the number of cameras. Our current system runs at several frames per second, and we believe that it is the first system capable of computing high-quality alpha mattes at near real-time rates without the use of active illumination or special backgrounds. Neel Joshi, Wojciech Matusik, Shai Avidan |
ACM Trans. Graph. | 3 |
| 2005 | Ensemble TrackingabstractWe consider tracking as a binary classification problem, where an ensemble of weak classifiers is trained online to distinguish between the object and the background. The ensemble of weak classifiers is combined into a strong classifier using AdaBoost. The strong classifier is then used to label pixels in the next frame as either belonging to the object or the background, giving a confidence map. The peak of the map, and hence the new position of the object, is found using mean shift. Temporal coherence is maintained by updating the ensemble with new weak classifiers that are trained online during tracking. We show a realization of this method and demonstrate it on several video sequences. Shai Avidan |
CVPR (2) | 1 |
| 2005 | Learning a Sparse, Corner-Based Representation for Time-varying Background ModelingabstractTime-varying phenomenon, such as ripples on water, trees waving in the wind and illumination changes, produces false motions, which significantly compromises the performance of an outdoor-surveillance system. In this paper, we propose a corner-based background model to effectively detect moving-objects in challenging dynamic scenes. Specifically, the method follows a three-step process. First, we detect feature points using a Harris corner detector and represent them as SIFT-like descriptors. Second, we dynamically learn a background model and classify each extracted feature as either a background or a foreground feature. Last, a "Lucas-Kanade" feature tracker is integrated into this framework to differentiate motion-consistent foreground objects from background objects with random or repetitive motion. The key insight of our work is that a collection of SIFT-like features can effectively represent the environment and account for variations caused by natural effects with dynamic movements. Features that do not correspond to the background must therefore correspond to foreground moving objects. Our method is computational efficient and works in real-time. Experiments on challenging video clips demonstrate that the proposed method achieves a higher accuracy in detecting the foreground objects than the existing methods. Qiang Zhu 0006, Shai Avidan, Kwang-Ting Cheng |
ICCV | 2 |
| 2005 | Spectral Bounds for Sparse PCA: Exact and Greedy AlgorithmsabstractSparse PCA seeks approximate sparse "eigenvectors" whose projections capture the maximal variance of data. As a cardinality-constrained and non-convex optimization problem, it is NP-hard and is encountered in a wide range of applied fields, from bio-informatics to finance. Recent progress has focused mainly on continuous approximation and convex relaxation of the hard cardinality constraint. In contrast, we consider an alternative discrete spectral formulation based on variational eigenvalue bounds and provide an effective greedy strategy as well as provably optimal solutions using branch-and-bound search. Moreover, the exact methodology used reveals a simple renormalization step that improves approximate solutions obtained by any continuous method. The resulting performance gain of discrete algorithms is demonstrated on real-world benchmark data and in extensive Monte Carlo evaluation trials. Baback Moghaddam, Yair Weiss, Shai Avidan |
NIPS | 3 |
| 2004 | Joint Feature-Basis Subset Selection
Shai Avidan |
CVPR (1) | 1 |
| 2004 | Probabilistic Multi-view Correspondence in a Distributed Setting with No Central Server
Shai Avidan, Yael Moses, Yoram Moses |
ECCV (4) | 1 |
| 2004 | The power of feature clustering: An application to object detectionabstractWe give a fast rejection scheme that is based on image segments and demonstrate it on the canonical example of face detection. However, in- stead of focusing on the detection step we focus on the rejection step and show that our method is simple and fast to be learned, thus making it an excellent pre-processing step to accelerate standard machine learning classifiers, such as neural-networks, Bayes classifiers or SVM. We de- compose a collection of face images into regions of pixels with similar behavior over the image set. The relationships between the mean and variance of image segments are used to form a cascade of rejectors that can reject over 99.8% of image patches, thus only a small fraction of the image patches must be passed to a full-scale classifier. Moreover, the training time for our method is much less than an hour, on a standard PC. The shape of the features (i.e. image segments) we use is data-driven, they are very cheap to compute and they form a very low dimensional feature space in which exhaustive search for the best features is tractable. Shai Avidan, Moshe Butman |
NIPS | 1 |
| 2004 | Support Vector TrackingabstractSupport Vector Tracking (SVT) integrates the Support Vector Machine (SVM) classifier into an optic-flow-based tracker. Instead of minimizing an intensity difference function between successive frames, SVT maximizes the SVM classification score. To account for large motions between successive frames, we build pyramids from the support vectors and use a coarse-to-fine approach in the classification stage. We show results of using SVT for vehicle tracking in image sequences. Shai Avidan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Subset Selection for Efficient SVM TrackingabstractWe update the SVM (support vector machine) score of an object through a video sequence with a small and variable subset of support vectors. In the first frame we use all the support vectors to compute the SVM score of the object but in subsequent frames we use only a small and variable subset of support vectors to update the SVM score. In each frame we calculate the dot-products of the support vectors in the subset with the pattern of the object being tracked. The difference in the dot-products, between past and current frames, is used to update the SVM score. This is done at a fraction of the computational cost required to re-evaluate the SVM score from scratch in every frame. The two methods we develop are "cyclic subset selection", in which we break the set of all support vectors into subsets of equal size and use them cyclically, and "maximum variance subset selection", in which we choose the support vectors whose dot-product with the test pattern varied the most in previous frames. We combine these techniques together for the problem of maintaining the SVM score of objects through a video sequence. Results on real video sequences are shown. Shai Avidan |
CVPR (1) | 1 |
| 2002 | EigenSegments: A Spatio-Temporal Decomposition of an Ensemble of Images
Shai Avidan |
ECCV (3) | 1 |
| 2001 | Support Vector TrackingabstractSupport Vector Tracking (SVT) integrates the Support Vector Machine (SVM) classifier into an optic-flow based tracker. Instead of minimizing an intensity difference function between successive frames, SVT maximizes the SVM classification score. To account for large motions between successive frames, we build pyramids from the support vectors and use a coarse-to-fine approach in the classification stage. We show results of using a homogeneous quadratic polynomial kernel-SVT for vehicle tracking in image sequences. Shai Avidan |
CVPR (1) | 1 |
| 2001 | Threading Fundamental MatricesabstractWe present a new function that operates on fundamental matrices across a sequence of views. The operation, we call "threading", connects two consecutive fundamental matrices using the trifocal tensor as the connecting thread. The threading operation guarantees that consecutive camera matrices are consistent with a unique 3D model, without ever recovering a 3D model. Applications include recovery of camera ego-motion from a sequence of views, image stabilization across a sequence, and multi-view image based rendering. Shai Avidan, Amnon Shashua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Layer Extraction from Multiple Images Containing Reflections and TransparencyabstractMany natural images contain reflections and transparency, i.e., they contain mixtures of reflected and transmitted light. When viewed from a moving camera, these appear as the superposition of component layer images moving relative to each other. The problem of multiple motion recovery has been previously studied by a number of researchers. However no one has yet demonstrated how to accurately recover the component images themselves. In this paper we develop an optimal approach to recovering layer images and their associated motions from an arbitrary number of composite images. We develop two different techniques for estimating the component layer images given known motion estimates. The first approach uses constrained least squares to recover the layer images. The second approach iteratively refines lower and upper bounds on the layer images using two novel compositing operations, namely minimum- and maximum-composites of aligned images. We combine these layer extraction techniques with a dominant motion estimator and a subsequent motion refinement stage. This results in a completely automated system that recovers transparent images and motions from a collection of input images. Richard Szeliski, Shai Avidan, P. Anandan 0001 |
CVPR | 2 |
| 2000 | Integrating Local Affine into Global Projective Images in the Joint Image Space
P. Anandan 0001, Shai Avidan |
ECCV (1) | 2 |
| 2000 | On the Reprojection of 3D and 2D Scenes Without Explicit Model Selection
Amnon Shashua, Shai Avidan |
ECCV (1) | 2 |
| 2000 | Trajectory Triangulation: 3D Reconstruction of Moving Points from a Monocular Image SequenceabstractWe consider the problem of reconstructing the 3D coordinates of a moving point seen from a monocular moving camera, i.e., to reconstruct moving objects from line-of-sight measurements only. The task is feasible only when some constraints are placed on the shape of the trajectory of the moving point. We coin the family of such tasks as "trajectory triangulation." We investigate the solutions for points moving along a straight-line and along conic-section trajectories, We show that if the point is moving along a straight line, then the parameters of the line (and, hence, the 3D position of the point at each time instant) can be uniquely recovered, and by linear methods, from at least five views. For the case of conic-shaped trajectory, we show that generally nine views are sufficient for a unique reconstruction of the moving point and fewer views when the conic is of a known type (like a circle in 3D Euclidean space for which seven views are sufficient). The paradigm of trajectory triangulation, in general, pushes the envelope of processing dynamic scenes forward. Thus static scenes become a particular case of a more general task of reconstructing scenes rich with moving objects (where an object could be a single point). Shai Avidan, Amnon Shashua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Trajectory Triangulation of Lines: Reconstruction of a 3D point Moving along a Line from a Monocular Image SequenceabstractWe consider the problem of reconstructing the location of a moving 3D point seen from a monocular moving camera, i.e., to reconstruct moving objects from line-of-sight measurements only. Since the point is moving while the camera is moving, then even if the camera motion is known, it is impossible to reconstruct the 3D location of the point under general circumstances. However we show that if the point is moving along a straight line, then the parameters of the line (and hence the 3D position of the point at each time instance) can be uniquely recovered, and by linear methods, from at least 5 views. Consequently, we propose a new approach for dealing with dynamic scenes (rich with moving objects) in which once the camera motion is recovered, the 3D trajectory (straight line) of the moving target can be recovered-even when the moving target consists of a single point. Shai Avidan, Amnon Shashua |
CVPR | 1 |
| 1999 | Trajectory Triangulation over Conic SectionsabstractWe consider the problem of reconstructing the 3D coordinates of a moving point seen from a monocular moving camera, i.e., to reconstruct moving objects from line-of-sight measurements only. The task is feasible only when same constraints are placed on the shape of the trajectory of the moving point. We coin the family of such tasks as "trajectory triangulation". In this paper we focus on trajectories whose shape is a conic-section and show that generally 9 views are sufficient for a unique reconstruction of the moving point and fewer views when the conic is a known type (like a circle in 3D Euclidean space for which 7 views are sufficient). Experiments demonstrate that our solutions are practical. The paradigm of Trajectory Triangulation in general pushes the envelope of processing dynamic scenes forward. Thus static scenes become a particular case of a more general task of reconstructing scenes rich with moving objects (where an object could be a single point). Amnon Shashua, Shai Avidan, Michael Werman |
ICCV | 2 |
| 1998 | Threading Fundamental Matrices
Shai Avidan, Amnon Shashua |
ECCV (1) | 1 |
| 1998 | Novel View Synthesis by Cascading Trilinear TensorsabstractPresents a new method for synthesizing novel views of a 3D scene from two or three reference images in full correspondence. The core of this work is the use and manipulation of an algebraic entity, termed the "trilinear tensor", that links point correspondences across three images. For a given virtual camera position and orientation, a new trilinear tensor can be computed based on the original tensor of the reference images. The desired view can then be created using this new trilinear tensor and point correspondences across two of the reference images. Shai Avidan, Amnon Shashua |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 1997 | Novel view synthesis in tensor spaceabstractWe present a new method for synthesizing novel views of a 3D scene from few model images in full correspondence. The core of this work is the derivation of a tensorial operator that describes the transformation from a given tensor of three views to a novel tensor of a new configuration of three views. By repeated application of the operator on a seed tensor with a sequence of desired virtual camera positions we obtain a chain of warping functions (tensors) from the set of model images to create the desired virtual views. Shai Avidan, Amnon Shashua |
CVPR | 1 |
| 1997 | Image-based view synthesis by combining trilinear tensors and learning techniquesabstractWe present a new method for rendering novel images of flexible 3D objects from a small number of example images in correspondence.The strength of the method is the ability to synthesize images whose viewing position is significantly far away from the viewing cone of the example images ("view extrapolation"), yet without ever modeling the 3D structure of the scene.The method relies on synthesizing a chain of "trilinear tensors" that govems the warping function from the example images to the novel image, together with a multi-dimensional interpolation function that synthesizes the non-rigidmotions of the viewed object from the virtual camera position.We show that two closely spaced example images alone are sufficient in practice to synthesize a significant viewing cone, thus demonstrating the ability of representing an object by a relatively small number of model images -for the purpose of cheap and fast viewers that can run on standard hardware. Shai Avidan, Theodoros Evgeniou, Amnon Shashua, Tomaso A. Poggio |
VRST | 1 |
| 1996 | Robust Recovery of Camera Rotation from Three FramesabstractComputing camera rotation from image sequences can be used for image stabilization, and when the camera rotation is known the computation of translation and scene structure are much simplified as well. A robust approach for recovering camera rotation is presented, which does not assume any specific scene structure (e.g. no planar surface is required), and which avoids prior computation of the epipole. Given two images taken from two different viewing positions, the rotation matrix between the images can be computed from any three homography matrices. The homographies are computed using the trilinear tensor which describes the relations between the projections of a 3D point into three images. The entire computation is linear for small angles, and is therefore fast and stable. Iterating the linear computation can then be used to recover larger rotations as well. Benny Rousso, Shai Avidan, Amnon Shashua, Shmuel Peleg |
CVPR | 2 |
| 1996 | The Rank 4 Constraint in Multiple (>=3) View Geometry
Amnon Shashua, Shai Avidan |
ECCV (2) | 2 |