Michal Irani

dblp:04/3190 · DBLP profile ↗
← Back
98ranked-venue papers
26as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 18 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 67 · 19 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers
abstract
Rerankers play a critical role in multimodal Retrieval-Augmented Generation (RAG) by refining ranking of an initial set of retrieved documents.Rerankers are typically trained using hard negative mining, whose goal is to select pages for each query which rank high, but are actually irrelevant.However, this selection process is typically passive and restricted to what the retriever can find in the available corpus, leading to several inherent limitations.These include: limited diversity, negative examples which are often not hard enough, low controllability, and frequent false negatives which harm training.Our paper proposes an alternative approach: Single-Page Hard Negative Query Generation, which goes the other way around.Instead of retrieving negative pages per query, we generate hard negative queries per page.Using an automated LLM-VLM pipeline, and given a page and its positive query, we create hard negatives by rephrasing the query to be as similar as possible in form and context, yet not answerable from the page.This paradigm enables fine-grained control over the generated queries, resulting in diverse, hard, and targeted negatives.It also supports efficient false negative verification.Our experiments show that rerankers trained with data generated using our approach outperform existing models and significantly improve retrieval performance 1 .
Navve Wasserman, Oliver Heinimann, Yuval Golbari, Tal Zimbalist, Eli Schwartz, Michal Irani
EMNLP6
2024 The Hidden Language of Diffusion Models
abstract
Text-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal representation of a textual concept by a diffusion model. This interpretation is obtained by decomposing the concept into a small set of human-interpretable textual elements. Applied over the state-of-the-art Stable Diffusion model, Conceptor reveals non-trivial structures in the representations of concepts. For example, we find surprising visual connections between concepts, that transcend their textual semantics. We additionally discover concepts that rely on mixtures of exemplars, biases, renowned artistic styles, or a simultaneous fusion of multiple meanings of the concept. Through a large battery of experiments, we demonstrate Conceptor's ability to provide meaningful, robust, and faithful decompositions for a wide variety of abstract, concrete, and complex textual concepts, while allowing to naturally connect each decomposition element to its corresponding visual impact on the generated images.
Hila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin, Assaf Shocher, Michal Irani, Inbar Mosseri, Lior Wolf
ICLR6
2023 Imagic: Text-Based Real Image Editing with Diffusion Models
abstract
Text-conditioned image editing has recently attracted considerable interest. However, most methods are currently limited to one of the following: specific editing types (e.g., object overlay, style transfer), synthetically generated images, or requiring multiple input images of a common object. In this paper we demonstrate, for the very first time, the ability to apply complex (e.g., non-rigid) text-based semantic edits to a single real image. For example, we can change the posture and composition of one or multiple objects inside an image, while preserving its original characteristics. Our method can make a standing dog sit down, cause a bird to spread its wings, etc. – each within its single high-resolution user-provided natural image. Contrary to previous work, our proposed method requires only a single input image and a target text (the desired edit). It operates on real images, and does not require any additional inputs (such as image masks or additional views of the object). Our method, called Imagic, leverages a pre-trained text-to-image diffusion model for this task. It produces a text embedding that aligns with both the input image and the target text, while fine-tuning the diffusion model to capture the image-specific appearance. We demonstrate the quality and versatility of Imagic on numerous inputs from various domains, showcasing a plethora of high quality complex semantic image edits, all within a single unified framework. To better assess performance, we introduce TEdBench, a highly challenging image editing benchmark. We conduct a user study, whose findings show that human raters prefer Imagic to previous leading editing methods on TEdBench.
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, Michal Irani
CVPR8
2023 Teaching CLIP to Count to Ten
abstract
Large vision-language models (VLMs), such as CLIP, learn rich joint image-text representations, facilitating advances in numerous downstream tasks, including zero-shot classification and text-to-image generation. Nevertheless, existing VLMs exhibit a prominent well-documented limitation – they fail to encapsulate compositional concepts such as counting. We introduce a simple yet effective method to improve the quantitative understanding of VLMs, while maintaining their overall performance on common benchmarks. Specifically, we propose a new counting-contrastive loss used to finetune a pre-trained VLM in tandem with its original objective. Our counting loss is deployed over automatically-created counterfactual examples, each consisting of an image and a caption containing an incorrect object count. For example, an image depicting three dogs is paired with the caption "Six dogs playing in the yard" as a negative example. Our loss encourages discrimination between the correct caption and its counterfactual variant which serves as a hard negative example. To the best of our knowledge, this work is the first to extend CLIP’s capabilities to object counting. Furthermore, we introduce "CountBench" – a new image-text counting benchmark for evaluating object counting capabilities. We demonstrate a significant improvement over state-of-the-art baseline models on this task. Finally, we leverage our counting-aware CLIP model for image retrieval and text-conditioned image generation, demonstrating that our model can produce specific counts of objects more reliably than existing ones.
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, Tali Dekel
ICCV6
2023 SinFusion: Training Diffusion Models on a Single Image or Video
abstract
Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image or video. In this paper we show how this can be resolved by training a diffusion model on a single input image or video. Our image/video-specific diffusion model (SinFusion) learns the appearance and dynamics of the single image or video, while utilizing the conditioning capabilities of diffusion models. It can solve a wide array of image/video-specific manipulation tasks. In particular, our model can learn from few frames the motion and dynamics of a single input video. It can then generate diverse new video samples of the same dynamic scene, extrapolate short videos into long ones (both forward and backward in time) and perform video upsampling. Most of these tasks are not realizable by current video-specific generation methods.
Yaniv Nikankin, Niv Haim, Michal Irani
ICML3
2023 Deconstructing Data Reconstruction: Multiclass, Weight Decay and General Losses
abstract
Memorization of training data is an active research area, yet our understanding of the inner workings of neural networks is still in its infancy. Recently, Haim et al. 2022 proposed a scheme to reconstruct training samples from multilayer perceptron binary classifiers, effectively demonstrating that a large portion of training samples are encoded in the parameters of such networks. In this work, we extend their findings in several directions, including reconstruction from multiclass and convolutional neural networks. We derive a more general reconstruction scheme which is applicable to a wider range of loss functions such as regression losses. Moreover, we study the various factors that contribute to networks' susceptibility to such reconstruction schemes. Intriguingly, we observe that using weight decay during training increases reconstructability both in terms of quantity and quality. Additionally, we examine the influence of the number of neurons relative to the number of training samples on the reconstructability. Code: https://github.com/gonbuzaglo/decoreco
Gon Buzaglo, Niv Haim, Gilad Yehudai, Gal Vardi, Yakir Oz, Yaniv Nikankin, Michal Irani
NeurIPS7
2022 Drop the GAN: In Defense of Patches Nearest Neighbors as Single Image Generative Models
abstract
Image manipulation dates back long before the deep learning era. The classical prevailing approaches were based on maximizing patch similarity between the input and generated output. Recently, single-image GANs were introduced as a superior and more sophisticated solution to image manipulation tasks. Moreover, they offered the opportunity not only to manipulate a given image, but also to generate a large and diverse set of different outputs from a single natural image. This gave rise to new tasks, which are considered “GAN-only”. However, despite their impressiveness, single-image GANs require long training time (usually hours) for each image and each task and often suffer from visual artifacts. In this paper we revisit the classical patch-based methods, and show that - unlike previously believed - classical methods can be adapted to tackle these novel “GAN-only” tasks. Moreover, they do so better and faster than single-image GAN-based methods. More specifically, we show that: (i) by introducing slight modifications, classical patch-based methods are able to unconditionally generate diverse images based on a single natural image; (ii) the generated output visual quality exceeds that of single-image GANs by a large margin (confirmed both quantitatively and qualitatively); (iii) they are orders of magnitude faster (runtime reduced from hours to seconds).22This project received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No 788535), and the Carolito Stiftung. Dr Bagon is a Robin Chemers Neustein AI Fellow.
Niv Granot, Ben Feinstein, Assaf Shocher, Shai Bagon, Michal Irani
CVPR5
2022 Diverse Generation from a Single Video Made Possible
Niv Haim, Ben Feinstein, Niv Granot, Assaf Shocher, Shai Bagon, Tali Dekel, Michal Irani
ECCV (17)7
2022 Combining Internal and External Constraints for Unrolling Shutter in Videos
Eyal Naor, Itai Antebi, Shai Bagon, Michal Irani
ECCV (17)4
2022 Pure Noise to the Rescue of Insufficient Data: Improving Imbalanced Classification by Training on Random Noise Images
abstract
Despite remarkable progress on visual recognition tasks, deep neural-nets still struggle to generalize well when training data is scarce or highly imbalanced, rendering them extremely vulnerable to real-world examples. In this paper, we present a surprisingly simple yet highly effective method to mitigate this limitation: using pure noise images as additional training data. Unlike the common use of additive noise or adversarial noise for data augmentation, we propose an entirely different perspective by directly training on pure random noise images. We present a new Distribution-Aware Routing Batch Normalization layer (DAR-BN), which enables training on pure noise images in addition to natural images within the same network. This encourages generalization and suppresses overfitting. Our proposed method significantly improves imbalanced classification performance, obtaining state-of-the-art results on a large variety of long-tailed image classification datasets (CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, Places-LT, and CelebA-5). Furthermore, our method is extremely simple and easy to use as a general new augmentation tool (on top of existing augmentations), and can be incorporated in any training scheme. It does not require any specialized data generation or training procedures, thus keeping training fast and efficient.
Shiran Zada, Itay Benou, Michal Irani
ICML3
2022 Reconstructing Training Data From Trained Neural Networks
abstract
Understanding to what extent neural networks memorize training data is an intriguing question with practical and theoretical implications. In this paper we show that in some cases a significant fraction of the training data can in fact be reconstructed from the parameters of a trained neural network classifier.We propose a novel reconstruction scheme that stems from recent theoretical results about the implicit bias in training neural networks with gradient-based methods.To the best of our knowledge, our results are the first to show that reconstructing a large portion of the actual training samples from a trained neural network classifier is generally possible.This has negative implications on privacy, as it can be used as an attack for revealing sensitive training data. We demonstrate our method for binary MLP classifiers on a few standard computer vision datasets.
Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, Michal Irani
NeurIPS5
2021 Explaining in Style: Training a GAN to explain a classifier in StyleSpace
abstract
Image classification models can depend on multiple different semantic attributes of the image. An explanation of the decision of the classifier needs to both discover and visualize these properties. Here we present StylEx, a method for doing this, by training a generative model to specifically explain multiple attributes that underlie classifier decisions. A natural source for such attributes is the StyleSpace of StyleGAN, which is known to generate semantically meaningful dimensions in the image. However, because standard GAN training is not dependent on the classifier, it may not represent those attributes which are important for the classifier decision, and the dimensions of StyleSpace may represent irrelevant at-tributes. To overcome this, we propose a training procedure for a StyleGAN, which incorporates the classifier model, in order to learn a classifier-specific StyleSpace. Explanatory attributes are then selected from this space. These can be used to visualize the effect of changing multiple attributes per image, thus providing image-specific explanations. We apply StylEx to multiple domains, including animals, leaves, faces and retinal images. For these, we show how an image can be modified in different ways to change its classifier output. Our results show that the method finds attributes that align well with semantic ones, generate meaningful image-specific explanations, and are human-interpretable as measured in user-studies.1
Oran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald, Gal Elidan, Avinatan Hassidim, William T. Freeman, Phillip Isola, Amir Globerson, Michal Irani, Inbar Mosseri
ICCV10
2020 SpeedNet: Learning the Speediness in Videos
abstract
We wish to automatically predict the “speediness” of moving objects in videos - whether they move faster, at, or slower than their “natural” speed. The core component in our approach is SpeedNet - a novel deep network trained to detect if a video is playing at normal rate, or if it is sped up. SpeedNet is trained on a large corpus of natural videos in a self-supervised manner, without requiring any manual annotations. We show how this single, binary classification network can be used to detect arbitrary rates of speediness of objects. We demonstrate prediction results by SpeedNet on a wide range of videos containing complex natural motions, and examine the visual cues it utilizes for making those predictions. Importantly, we show that through predicting the speed of videos, the model learns a powerful and meaningful space-time representation that goes beyond simple motion cues. We demonstrate how those learned features can boost the performance of self supervised action recognition, and can be used for video retrieval. Furthermore, we also apply SpeedNet for generating time-varying, adaptive video speedups, which can allow viewers to watch videos faster, but with less of the jittery, unnatural motions typical to videos that are sped up uniformly.
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, Tali Dekel
CVPR7
2020 Semantic Pyramid for Image Generation
abstract
We present a novel GAN-based model that utilizes the space of deep features learned by a pre-trained classification model. Inspired by classical image pyramid representations, we construct our model as a Semantic Generation Pyramid -- a hierarchical framework which leverages the continuum of semantic information encapsulated in such deep features; this ranges from low level information contained in fine features to high level, semantic information contained in deeper features. More specifically, given a set of features extracted from a reference image, our model generates diverse image samples, each with matching features at each semantic level of the classification model. We demonstrate that our model results in a versatile and flexible framework that can be used in various classic and novel image generation tasks. These include: generating images with a controllable extent of semantic similarity to a reference image, and different manipulation tasks such as semantically-controlled inpainting and compositing; all achieved with the same model, with no further training.
Assaf Shocher, Yossi Gandelsman, Inbar Mosseri, Michal Yarom, Michal Irani, William T. Freeman, Tali Dekel
CVPR5
2020 Across Scales and Across Dimensions: Temporal Super-Resolution Using Deep Internal Learning
Liad Pollak Zuckerman, Eyal Naor, George Pisha, Shai Bagon, Michal Irani
ECCV (7)5
2019 "Double-DIP": Unsupervised Image Decomposition via Coupled Deep-Image-Priors
abstract
Many seemingly unrelated computer vision tasks can be viewed as a special case of image decomposition into separate layers. For example, image segmentation (separation into foreground and background layers); transparent layer separation (into reflection and transmission layers); Image dehazing (separation into a clear image and a haze map), and more. In this paper we propose a unified framework for unsupervised layer decomposition of a single image, based on coupled "Deep-image-Prior" (DIP) networks. It was shown [Ulyanov et al] that the structure of a single DIP generator network is sufficient to capture the low-level statistics of a single image. We show that coupling multiple such DIPs provides a powerful tool for decomposing images into their basic components, for a wide variety of applications. This capability stems from the fact that the internal statistics of a mixture of layers is more complex than the statistics of each of its individual components. We show the power of this approach for Image-Dehazing, Fg/Bg Segmentation, Watermark-Removal, Transparency Separation in images and video, and more. These capabilities are achieved in a totally unsupervised way, with no training examples other than the input image/video itself.
Yossi Gandelsman, Assaf Shocher, Michal Irani
CVPR3
2019 InGAN: Capturing and Retargeting the "DNA" of a Natural Image
abstract
Generative Adversarial Networks (GANs) typically learn a distribution of images in a large image dataset, and are then able to generate new images from this distribution. However, each natural image has its own internal statistics, captured by its unique distribution of patches. In this paper we propose an "Internal GAN'' (InGAN) - an image-specific GAN - which trains on a single input image and learns its internal distribution of patches. It is then able to synthesize a plethora of new natural images of significantly different sizes, shapes and aspect-ratios - all with the same internal patch-distribution (same "DNA'') as the input image. In particular, despite large changes in global size/shape of the image, all elements inside the image maintain their local size/shape. InGAN is fully unsupervised, requiring no additional data other than the input image itself. Once trained on the input image, it can remap the input to any size or shape in a single feedforward pass, while preserving the same internal patch distribution. InGAN provides a unified framework for a variety of tasks, bridging the gap between textures and natural images.
Assaf Shocher, Shai Bagon, Phillip Isola, Michal Irani
ICCV4
2019 Unsupervised Internal Learning
Michal Irani
ICPRAM1
2019 From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI
abstract
Reconstructing observed images from fMRI brain recordings is challenging. Unfortunately, acquiring sufficient ''labeled'' pairs of {Image, fMRI} (i.e., images with their corresponding fMRI responses) to span the huge space of natural images is prohibitive for many reasons. We present a novel approach which, in addition to the scarce labeled data (training pairs), allows to train fMRI-to-image reconstruction networks also on "unlabeled" data (i.e., images without fMRI recording, and fMRI recording without images). The proposed model utilizes both an Encoder network (image-to-fMRI) and a Decoder network (fMRI-to-image). Concatenating these two networks back-to-back (Encoder-Decoder & Decoder-Encoder) allows augmenting the training data with both types of unlabeled data. Importantly, it allows training on the unlabeled test-fMRI data. This self-supervision adapts the reconstruction network to the new input test-data, despite its deviation from the statistics of the scarce training data.
Roman Beliy, Guy Gaziv, Assaf Hoogi, Francesca Strappini, Tal Golan, Michal Irani
NeurIPS6
2019 Blind Super-Resolution Kernel Estimation using an Internal-GAN
abstract
Super resolution (SR) methods typically assume that the low-resolution (LR) image was downscaled from the unknown high-resolution (HR) image by a fixed `ideal’ downscaling kernel (e.g. Bicubic downscaling). However, this is rarely the case in real LR images, in contrast to synthetically generated SR datasets. When the assumed downscaling kernel deviates from the true one, the performance of SR methods significantly deteriorates. This gave rise to Blind-SR - namely, SR when the downscaling kernel (SR-kernel’’) is unknown. It was further shown that the true SR-kernel is the one that maximizes the recurrence of patches across scales of the LR image. In this paper we show how this powerful cross-scale recurrence property can be realized using Deep Internal Learning. We introduceKernelGAN’’, an image-specific Internal-GAN, which trains solely on the LR test image at test time, and learns its internal distribution of patches. Its Generator is trained to produce a downscaled version of the LR test image, such that its Discriminator cannot distinguish between the patch distribution of the downscaled image, and the patch distribution of the original LR image. The Generator, once trained, constitutes the downscaling operation with the correct image-specific SR-kernel. KernelGAN is fully unsupervised, requires no training data other than the input image itself, and leads to state-of-the-art results in Blind-SR when plugged into existing SR algorithms.
Sefi Bell-Kligler, Assaf Shocher, Michal Irani
NeurIPS3
2019 "Blind" visual inference by composition
Michal Irani
Pattern Recognit. Lett.1
2018 "Zero-Shot" Super-Resolution Using Deep Internal Learning
abstract
Deep Learning has led to a dramatic leap in SuperResolution (SR) performance in the past few years. However, being supervised, these SR methods are restricted to specific training data, where the acquisition of the low-resolution (LR) images from their high-resolution (HR) counterparts is predetermined (e.g., bicubic downscaling), without any distracting artifacts (e.g., sensor noise, image compression, non-ideal PSF, etc). Real LR images, however, rarely obey these restrictions, resulting in poor SR results by SotA (State of the Art) methods. In this paper we introduce "Zero-Shot" SR, which exploits the power of Deep Learning, but does not rely on prior training. We exploit the internal recurrence of information inside a single image, and train a small image-specific CNN at test time, on examples extracted solely from the input image itself. As such, it can adapt itself to different settings per image. This allows to perform SR of real old photos, noisy images, biological data, and other images where the acquisition process is unknown or non-ideal. On such images, our method outperforms SotA CNN-based SR methods, as well as previous unsupervised SR methods. To the best of our knowledge, this is the first unsupervised CNN-based SR method.
Assaf Shocher, Michal Irani
CVPR3
2017 Non-uniform Blind Deblurring by Reblurring
abstract
We present an approach for blind image deblurring, which handles non-uniform blurs. Our algorithm has two main components: (i) A new method for recovering the unknown blur-field directly from the blurry image, and (ii) A method for deblurring the image given the recovered non-uniform blur-field. Our blur-field estimation is based on analyzing the spectral content of blurry image patches by Re-blurring them. Being unrestricted by any training data, it can handle a large variety of blur sizes, yielding superior blur-field estimation results compared to training-based deep-learning methods. Our non-uniform deblurring algorithm is based on the internal image-specific patch-recurrence prior. It attempts to recover a sharp image which, on one hand - results in the blurry image under our estimated blur-field, and on the other hand - maximizes the internal recurrence of patches within and across scales of the recovered sharp image. The combination of these two components gives rise to a blind-deblurring algorithm, which exceeds the performance of state-of-the-art CNN-based blind-deblurring by a significant margin, without the need for any training data.
Yuval Bahat, Netalee Efrat, Michal Irani
ICCV3
2016 Needle-Match: Reliable Patch Matching under High Uncertainty
abstract
Reliable patch-matching forms the basis for many algorithms (super-resolution, denoising, inpainting, etc.) However, when the image quality deteriorates (by noise, blur or geometric distortions), the reliability of patch-matching deteriorates as well. Matched patches in the degraded image, do not necessarily imply similarity of the underlying patches in the (unknown) high-quality image. This restricts the applicability of patch-based methods. In this paper we present a patch representation called "Needle", which consists of small multi-scale versions of the patch and its immediate surrounding region. While the patch at the finest image scale is severely degraded, the degradation decreases dramatically in coarser needle scales, revealing reliable information for matching. We show that the Needle is robust to many types of image degradations, leads to matches faithful to the underlying high-quality patches, and to improvement in existing patch-based methods.
Or Lotan, Michal Irani
CVPR2
2016 Blind dehazing using internal patch recurrence
abstract
Images of outdoor scenes are often degraded by haze, fog and other scattering phenomena. In this paper we show how such images can be dehazed using internal patch recurrence. Small image patches tend to repeat abundantly inside a natural image, both within the same scale, as well as across different scales. This behavior has been used as a strong prior for image denoising, super-resolution, image completion and more. Nevertheless, this strong recurrence property significantly diminishes when the imaging conditions are not ideal, as is the case in images taken under bad weather conditions (haze, fog, underwater scattering, etc.). In this paper we show how we can exploit the deviations from the ideal patch recurrence for "Blind De-hazing" — namely, recovering the unknown haze parameters and reconstructing a haze-free image. We seek the haze parameters that, when used for dehazing the input image, will maximize the patch recurrence in the dehazed output image. More specifically, pairs of co-occurring patches at different depths (hence undergoing different degrees of haze) allow recovery of the airlight color, as well as the relative-transmission of each such pair of patches. This in turn leads to dense recovery of the scene structure, and to full image dehazing.
Yuval Bahat, Michal Irani
ICCP2
2015 Revealing and modifying non-local variations in a single image
abstract
We present an algorithm for automatically detecting and visualizing small non-local variations between repeating structures in a single image. Our method allows to automatically correct these variations, thus producing an 'idealized' version of the image in which the resemblance between recurring structures is stronger. Alternatively, it can be used to magnify these variations, thus producing an exaggerated image which highlights the various variations that are difficult to spot in the input image. We formulate the estimation of deviations from perfect recurrence as a general optimization problem, and demonstrate it in the particular cases of geometric deformations and color variations.
Tali Dekel, Tomer Michaeli, Michal Irani, William T. Freeman
ACM Trans. Graph.3
2014 Video Segmentation by Non-Local Consensus voting
Alon Faktor, Michal Irani
BMVC2
2014 Blind Deblurring Using Internal Patch Recurrence
Tomer Michaeli, Michal Irani
ECCV (3)2
2014 "Clustering by Composition" - Unsupervised Discovery of Image Categories
abstract
We define a "good image cluster" as one in which images can be easily composed (like a puzzle) using pieces from each other, while are difficult to compose from images outside the cluster. The larger and more statistically significant the pieces are, the stronger the affinity between the images. This gives rise to unsupervised discovery of very challenging image categories. We further show how multiple images can be composed from each other simultaneously and efficiently using a collaborative randomized search algorithm. This collaborative process exploits the "wisdom of crowds of images", to obtain a sparse yet meaningful set of image affinities, and in time which is almost linear in the size of the image collection. "Clustering-by-Composition" yields state-of-the-art results on current benchmark data sets. It further yields promising results on new challenging data sets, such as data sets with very few images (where a `cluster model' cannot be `learned' by current methods), and a subset of the PASCAL VOC data set (with huge variability in scale and appearance).
Alon Faktor, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Separating Signal from Noise Using Patch Recurrence across Scales
abstract
Recurrence of small clean image patches across different scales of a natural image has been successfully used for solving ill-posed problems in clean images (e.g., super-resolution from a single image). In this paper we show how this multi-scale property can be extended to solve ill-posed problems under noisy conditions, such as image denoising. While clean patches are obscured by severe noise in the original scale of a noisy image, noise levels drop dramatically at coarser image scales. This allows for the unknown hidden clean patches to "naturally emerge" in some coarser scale of the noisy image. We further show that patch recurrence across scales is strengthened when using directional pyramids (that blur and sub sample only in one direction). Our statistical experiments show that for almost any noisy image patch (more than 99%), there exists a "good" clean version of itself at the same relative image coordinates in some coarser scale of the image. This is a strong phenomenon of noise-contaminated natural images, which can serve as a strong prior for separating the signal from the noise. Finally, incorporating this multi-scale prior into a simple denoising algorithm yields state-of-the-art denoising results.
Maria Zontak, Inbar Mosseri, Michal Irani
CVPR3
2013 Combining the power of Internal and External denoising
abstract
Image denoising methods can broadly be classified into two types: “Internal Denoising” (denoising an image patch using other noisy patches within the noisy image), and “External Denoising” (denoising a patch using external clean natural image patches). Any such method, whether Internal or External, is typically applied to all image patches. In this paper we show that different image patches inherently have different preferences for Internal or External de-noising. Moreover, and surprisingly, the higher the noise in the image, the stronger the preference for Internal De-noising. We identify and explain the source of this behavior, and show that Internal/External preference of a patch is directly related to its individual Signal-to-Noise-Ratio (“PatchSNR”). Patches with high PatchSNR (e.g., patches on strong edges) benefit much from External Denoising, whereas patches with low PatchSNR (e.g., patches in noisy uniform regions) benefit much more from Internal Denoising. Combining the power of Internal or External denoising selectively for each patch based on its estimated PatchSNR leads to improvement in denoising performance.
Inbar Mosseri, Maria Zontak, Michal Irani
ICCP3
2013 Co-segmentation by Composition
abstract
Given a set of images which share an object from the same semantic category, we would like to co-segment the shared object. We define 'good' co-segments to be ones which can be easily composed (like a puzzle) from large pieces of other co-segments, yet are difficult to compose from remaining image parts. These pieces must not only match well but also be statistically significant (hard to compose at random). This gives rise to co-segmentation of objects in very challenging scenarios with large variations in appearance, shape and large amounts of clutter. We further show how multiple images can collaborate and "score" each others' co-segments to improve the overall fidelity and accuracy of the co-segmentation. Our co-segmentation can be applied both to large image collections, as well as to very few images (where there is too little data for unsupervised learning). At the extreme, it can be applied even to a single image, to extract its co-occurring objects. Our approach obtains state-of-the-art results on benchmark datasets. We further show very encouraging co-segmentation results on the challenging PASCAL-VOC dataset.
Alon Faktor, Michal Irani
ICCV2
2013 Nonparametric Blind Super-resolution
abstract
Super resolution (SR) algorithms typically assume that the blur kernel is known (either the Point Spread Function 'PSF' of the camera, or some default low-pass filter, e.g. a Gaussian). However, the performance of SR methods significantly deteriorates when the assumed blur kernel deviates from the true one. We propose a general framework for "blind" super resolution. In particular, we show that: (i) Unlike the common belief, the PSF of the camera is the wrong blur kernel to use in SR algorithms. (ii) We show how the correct SR blur kernel can be recovered directly from the low-resolution image. This is done by exploiting the inherent recurrence property of small natural image patches (either internally within the same image, or externally in a collection of other natural images). In particular, we show that recurrence of small patches across scales of the low-res image (which forms the basis for single-image SR), can also be used for estimating the optimal blur kernel. This leads to significant improvement in SR results.
Tomer Michaeli, Michal Irani
ICCV2
2012 "Clustering by Composition" - Unsupervised Discovery of Image Categories
Alon Faktor, Michal Irani
ECCV (7)2
2011 Space-time super-resolution from a single video
abstract
Spatial Super Resolution (SR) aims to recover fine image details, smaller than a pixel size. Temporal SR aims to recover rapid dynamic events that occur faster than the video frame-rate, and are therefore invisible or seen incorrectly in the video sequence. Previous methods for Space-Time SR combined information from multiple video recordings of the same dynamic scene. In this paper we show how this can be done from a single video recording. Our approach is based on the observation that small space-time patches (`ST-patches', e.g., 5×5×3) of a single `natural video', recur many times inside the same video sequence at multiple spatio-temporal scales. We statistically explore the degree of these ST-patch recurrences inside `natural videos', and show that this is a very strong statistical phenomenon. Space-time SR is obtained by combining information from multiple ST-patches at sub-frame accuracy. We show how finding similar ST-patches can be done both efficiently (with a randomized-based search in space-time), and at sub-frame accuracy (despite severe motion aliasing). Our approach is particularly useful for temporal SR, resolving both severe motion aliasing and severe motion blur in complex `natural videos'.
Oded Shahar, Alon Faktor, Michal Irani
CVPR3
2011 Internal statistics of a single natural image
abstract
Statistics of `natural images' provides useful priors for solving under-constrained problems in Computer Vision. Such statistics is usually obtained from large collections of natural images. We claim that the substantial internal data redundancy within a single natural image (e.g., recurrence of small image patches), gives rise to powerful internal statistics, obtained directly from the image itself. While internal patch recurrence has been used in various applications, we provide a parametric quantification of this property. We show that the likelihood of an image patch to recur at another image location can be expressed parametricly as a function of the spatial distance from the patch, and its gradient content. This “internal parametric prior” is used to improve existing algorithms that rely on patch recurrence. Moreover, we show that internal image-specific statistics is often more powerful than general external statistics, giving rise to more powerful image-specific priors. In particular: (i) Patches tend to recur much more frequently (densely) inside the same image, than in any random external collection of natural images. (ii) To find an equally good external representative patch for all the patches of an image, requires an external database of hundreds of natural images. (iii) Internal statistics often has stronger predictive power than external statistics, indicating that it may potentially give rise to more powerful image-specific priors.
Maria Zontak, Michal Irani
CVPR2
2010 Detecting and sketching the common
abstract
Given very few images containing a common object of interest under severe variations in appearance, we detect the common object and provide a compact visual representation of that object, depicted by a binary sketch. Our algorithm is composed of two stages: (i) Detect a mutually common (yet non-trivial) ensemble of `self-similarity descriptors' shared by all the input images. (ii) Having found such a mutually common ensemble, `invert' it to generate a compact sketch which best represents this ensemble. This provides a simple and compact visual representation of the common object, while eliminating the background clutter of the query images. It can be obtained from very few query images. Such clean sketches may be useful for detection, retrieval, recognition, co-segmentation, and for artistic graphical purposes.
Shai Bagon, Or Brostovski, Meirav Galun, Michal Irani
CVPR4
2010 Regenerative morphing
abstract
We present a new image morphing approach in which the output sequence is regenerated from small pieces of the two source (input) images. The approach does not require manual correspondence, and generates compelling results even when the images are of very different objects (e.g., a cloud and a face). We pose the morphing task as an optimization with the objective of achieving bidirectional similarity of each frame to its neighbors, and also to the source images. The advantages of this approach are 1) it can operate fully automatically, producing effective results for many sequences (but also supports manual correspondences, when available), 2) ghosting artifacts are minimized, and 3) different parts of the scene move at different rates, yielding more interesting (and less robotic) transitions.
Eli Shechtman, Alex Rav-Acha, Michal Irani, Steven M. Seitz
CVPR3
2009 Super-resolution from a single image
abstract
Methods for super-resolution can be broadly classified into two families of methods: (i) The classical multi-image super-resolution (combining images obtained at subpixel misalignments), and (ii) Example-Based super-resolution (learning correspondence between low and high resolution image patches from a database). In this paper we propose a unified framework for combining these two families of methods. We further show how this combined approach can be applied to obtain super resolution from as little as a single image (with no database or prior examples). Our approach is based on the observation that patches in a natural image tend to redundantly recur many times inside the image, both within the same scale, as well as across different scales. Recurrence of patches within the same image scale (at subpixel misalignments) gives rise to the classical super-resolution, whereas recurrence of patches across different scales of the same image gives rise to example-based super-resolution. Our approach attempts to recover at each pixel its best possible resolution increase based on its patch redundancy within and across scales.
Daniel Glasner, Shai Bagon, Michal Irani
ICCV3
2008 In defense of Nearest-Neighbor based image classification
abstract
State-of-the-art image classification methods require an intensive learning/training stage (using SVM, Boosting, etc.) In contrast, non-parametric nearest-neighbor (NN) based image classifiers require no training time and have other favorable properties. However, the large performance gap between these two families of approaches rendered NN-based image classifiers useless. We claim that the effectiveness of non-parametric NN-based image classification has been considerably undervalued. We argue that two practices commonly used in image classification methods, have led to the inferior performance of NN-based image classifiers: (i) Quantization of local image descriptors (used to generate "bags-of-words ", codebooks). (ii) Computation of 'image-to-image' distance, instead of 'image-to-class' distance. We propose a trivial NN-based classifier - NBNN, (Naive-Bayes nearest-neighbor), which employs NN- distances in the space of the local image descriptors (and not in the space of images). NBNN computes direct 'image- to-class' distances without descriptor quantization. We further show that under the Naive-Bayes assumption, the theoretically optimal image classifier can be accurately approximated by NBNN. Although NBNN is extremely simple, efficient, and requires no learning/training phase, its performance ranks among the top leading learning-based image classifiers. Empirical comparisons are shown on several challenging databases (Caltech-101 ,Caltech-256 and Graz-01).
Oren Boiman, Eli Shechtman, Michal Irani
CVPR3
2008 Summarizing visual data using bidirectional similarity
abstract
We propose a principled approach to summarization of visual data (images or video) based on optimization of a well-defined similarity measure. The problem we consider is re-targeting (or summarization) of image/video data into smaller sizes. A good ldquovisual summaryrdquo should satisfy two properties: (1) it should contain as much as possible visual information from the input data; (2) it should introduce as few as possible new visual artifacts that were not in the input data (i.e., preserve visual coherence). We propose a bi-directional similarity measure which quantitatively captures these two requirements: Two signals S and T are considered visually similar if all patches of S (at multiple scales) are contained in T, and vice versa. The problem of summarization/re-targeting is posed as an optimization problem of this bi-directional similarity measure. We show summarization results for image and video data. We further show that the same approach can be used to address a variety of other problems, including automatic cropping, completion and synthesis of visual data, image collage, object removal, photo reshuffling and more.
Denis Simakov, Yaron Caspi, Eli Shechtman, Michal Irani
CVPR4
2008 What Is a Good Image Segment? A Unified Approach to Segment Extraction
Shai Bagon, Oren Boiman, Michal Irani
ECCV (4)3
2007 Matching Local Self-Similarities across Images and Videos
abstract
We present an approach for measuring similarity between visual entities (images or videos) based on matching internal self-similarities. What is correlated across images (or across video sequences) is the internal layout of local self-similarities (up to some distortions), even though the patterns generating those local self-similarities are quite different in each of the images/videos. These internal self-similarities are efficiently captured by a compact local "self-similarity descriptor"', measured densely throughout the image/video, at multiple scales, while accounting for local and global geometric distortions. This gives rise to matching capabilities of complex visual data, including detection of objects in real cluttered images using only rough hand-sketches, handling textured objects with no clear boundaries, and detecting complex actions in cluttered video data with no prior learning. We compare our measure to commonly used image-based and video-based similarity measures, and demonstrate its applicability to object detection, retrieval, and action detection.
Eli Shechtman, Michal Irani
CVPR2
2007 Detecting Irregularities in Images and in Video
Oren Boiman, Michal Irani
Int. J. Comput. Vis.2
2007 Actions as Space-Time Shapes
abstract
Human action in video sequences can be seen as silhouettes of a moving torso and protruding limbs undergoing articulated motion. We regard human actions as three-dimensional shapes induced by the silhouettes in the space-time volume. We adopt a recent approach for analyzing 2D shapes and generalize it to deal with volumetric space-time action shapes. Our method utilizes properties of the solution to the Poisson equation to extract space-time features such as local space-time saliency, action dynamics, shape structure and orientation. We show that these features are useful for action recognition, detection and clustering. The method is fast, does not require video alignment and is applicable in (but not limited to) many scenarios where the background is known. Moreover, we demonstrate the robustness of our method to partial occlusions, non-rigid deformations, significant changes in scale and viewpoint, high irregularities in the performance of an action, and low quality video.
Lena Gorelick, Moshe Blank, Eli Shechtman, Michal Irani, Ronen Basri
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Space-Time Behavior-Based Correlation - OR - How to Tell If Two Underlying Motion Fields Are Similar Without Computing Them?
abstract
We introduce a behavior-based similarity measure which tells us whether two different space-time intensity patterns of two different video segments could have resulted from a similar underlying motion field. This is done directly from the intensity information, without explicitly computing the underlying motions. Such a measure allows us to detect similarity between video segments of differently dressed people performing the same type of activity. It requires no foreground/background segmentation, no prior learning of activities, and no motion estimation or tracking. Using this behavior-based similarity measure, we extend the notion of 2-dimensional image correlation into the 3-dimensional space-time volume, thus allowing to correlate dynamic behaviors and actions. Small space-time video segments (small video clips) are "correlated" against entire video sequences in all three dimensions (x,y, and t). Peak correlation values correspond to video locations with similar dynamic behaviors. Our approach can detect very complex behaviors in video sequences (e.g., ballet movements, pool dives, running water), even when multiple complex activities occur simultaneously within the field-of-view of the camera. We further show its robustness to small changes in scale and orientation of the correlated behavior.
Eli Shechtman, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Space-Time Completion of Video
abstract
This paper presents a new framework for the completion of missing information based on local structures. It poses the task of completion as a global optimization problem with a well-defined objective function and derives a new algorithm to optimize it. Missing values are constrained to form coherent structures with respect to reference examples. We apply this method to space-time completion of large space-time "holes" in video sequences of complex dynamic scenes. The missing portions are filled in by sampling spatio-temporal patches from the available parts of the video, while enforcing global spatio-temporal consistency between all patches in and around the hole. The consistent completion of static scene parts simultaneously with dynamic behaviors leads to realistic looking video sequences and images. Space-time video completion is useful for a variety of tasks, including, but not limited to: 1) Sophisticated video removal (of undesired static or dynamic objects) by completing the appropriate static or dynamic background information. 2) Correction of missing/corrupted video frames in old movies. 3) Modifying a visual story by replacing unwanted elements. 4) Creation of video textures by extending smaller ones. 5) Creation of complete field-of-view stabilized video. 6) As images are one-frame videos, we apply the method to this special case as well.
Yonatan Wexler, Eli Shechtman, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Aligning Sequences and Actions by Maximizing Space-Time Correlations
Yaron Ukrainitz, Michal Irani
ECCV (3)2
2006 Similarity by Composition
abstract
We propose a new approach for measuring similarity between two signals, which is applicable to many machine learning tasks, and to many signal types. We say that a signal S1 is “similar” to a signal S2 if it is “easy” to compose S1 from few large contiguous chunks of S2. Obviously, if we use small enough pieces, then any signal can be composed of any other. Therefore, the larger those pieces are, the more similar S1 is to S2. This induces a local similarity score at every point in the signal, based on the size of its supported surrounding region. These local scores can in turn be accumulated in a principled information-theoretic way into a global similarity score of the entire S1 to S2. “Similarity by Composition” can be applied between pairs of signals, between groups of signals, and also between dif- ferent portions of the same signal. It can therefore be employed in a wide variety of machine learning problems (clustering, classification, retrieval, segmentation, attention, saliency, labelling, etc.), and can be applied to a wide range of signal types (images, video, audio, biological data, etc.) We show a few such examples.
Oren Boiman, Michal Irani
NIPS2
2006 Feature-Based Sequence-to-Sequence Matching
Yaron Caspi, Denis Simakov, Michal Irani
Int. J. Comput. Vis.3
2006 On Single-Sequence and Multi-Sequence Factorizations
Lihi Zelnik-Manor, Michal Irani
Int. J. Comput. Vis.2
2006 Multi-body Factorization with Uncertainty: Revisiting Motion Consistency
Lihi Zelnik-Manor, Moshe Machline, Michal Irani
Int. J. Comput. Vis.3
2006 Statistical Analysis of Dynamic Actions
abstract
Real-world action recognition applications require the development of systems which are fast, can handle a large variety of actions without a priori knowledge of the type of actions, need a minimal number of parameters, and necessitate as short as possible learning stage. In this paper, we suggest such an approach. We regard dynamic activities as long-term temporal objects, which are characterized by spatio-temporal features at multiple temporal scales. Based on this, we design a simple statistical distance measure between video sequences which captures the similarities in their behavioral content. This measure is nonparametric and can thus handle a wide range of complex dynamic actions. Having a behavior-based distance measure between sequences, we use it for a variety of tasks, including: video indexing, temporal segmentation, and action-based video clustering. These tasks are performed without prior knowledge of the types of actions, their models, or their temporal extents.
Lihi Zelnik-Manor, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Space-Time Behavior Based Correlation
abstract
We introduce a behavior-based similarity measure which tells us whether two different space-time intensity patterns of two different video segments could have resulted from a similar underlying motion field. This is done directly from the intensity information, without explicitly computing the underlying motions. Such a measure allows us to detect similarity between video segments of differently dressed people performing the same type of activity. It requires no foreground/background segmentation, no prior learning of activities, and no motion estimation or tracking. Using this behavior-based similarity measure, we extend the notion of 2-dimensional image correlation into the 3-dimensional space-time volume, thus allowing to correlate dynamic behaviors and actions. Small space-time video segments (small video clips) are "correlated" against entire video sequences in all three dimensions (x,y, and t). Peak correlation values correspond to video locations with similar dynamic behaviors. Our approach can detect very complex behaviors in video sequences (e.g., ballet movements, pool dives, running water), even when multiple complex activities occur simultaneously within the field-of-view of the camera.
Eli Shechtman, Michal Irani
CVPR (1)2
2005 Actions as Space-Time Shapes
abstract
Human action in video sequences can be seen as silhouettes of a moving torso and protruding limbs undergoing articulated motion. We regard human actions as three-dimensional shapes induced by the silhouettes in the space-time volume. We adopt a recent approach by Gorelick et al. (2004) for analyzing 2D shapes and generalize it to deal with volumetric space-time action shapes. Our method utilizes properties of the solution to the Poisson equation to extract space-time features such as local space-time saliency, action dynamics, shape structure and orientation. We show that these features are useful for action recognition, detection and clustering. The method is fast, does not require video alignment and is applicable in (but not limited to) many scenarios where the background is known. Moreover, we demonstrate the robustness of our method to partial occlusions, non-rigid deformations, significant changes in scale and viewpoint, high irregularities in the performance of an action and low quality video
Moshe Blank, Lena Gorelick, Eli Shechtman, Michal Irani, Ronen Basri
ICCV4
2005 Detecting Irregularities in Images and in Video
abstract
We address the problem of detecting irregularities in visual data, e.g., detecting suspicious behaviors in video sequences, or identifying salient patterns in images. The term "irregular" depends on the context in which the "regular" or "valid" are defined. Yet, it is not realistic to expect explicit definition of all possible valid configurations for a given context. We pose the problem of determining the validity of visual data as a process of constructing a puzzle: We try to compose a new observed image region or a new video segment ("the query") using chunks of data ("pieces of puzzle") extracted from previous visual examples ("the database "). Regions in the observed data which can be composed using large contiguous chunks of data from the database are considered very likely, whereas regions in the observed data which cannot be composed from the database (or can be composed, but only using small fragmented pieces) are regarded as unlikely/suspicious. The problem is posed as an inference process in a probabilistic graphical model. We show applications of this approach to identifying saliency in images and video, and for suspicious behavior recognition.
Oren Boiman, Michal Irani
ICCV2
2005 Separating Transparent Layers of Repetitive Dynamic Behaviors
abstract
In this paper, we present an approach for separating two transparent layers of complex nonrigid scene dynamics. The dynamics in one of the layers is assumed to be repetitive, while the other can have any arbitrary dynamics. Such repetitive dynamics includes, among other, human actions in video (e.g., a walking person), or a repetitive musical tune in audio signals. We use a global to local space time alignment approach to detect and align the repetitive behavior. Once aligned, a median operator applied to space time derivatives is used to recover the intrinsic repeating behavior, and separate it from the other transparent layer. We show results on synthetic and real video sequences. In addition, we show the applicability of our approach to separating mixed audio signals (from a single source).
Bernard Sarel, Michal Irani
ICCV2
2005 Space-Time Super-Resolution
abstract
We propose a method for constructing a video sequence of high space-time resolution by combining information from multiple low-resolution video sequences of the same dynamic scene. Super-resolution is performed simultaneously in time and in space. By "temporal super-resolution," we mean recovering rapid dynamic events that occur faster than regular frame-rate. Such dynamic events are not visible (or else are observed incorrectly) in any of the input sequences, even if these are played in "slow-motion." The spatial and temporal dimensions are very different in nature, yet are interrelated. This leads to interesting visual trade-offs in time and space and to new video applications. These include: 1) treatment of spatial artifacts (e.g., motion-blur) by increasing the temporal resolution and 2) combination of input sequences of different space-time resolutions (e.g., NTSC, PAL, and even high quality still images) to generate a high quality video sequence. We further analyze and compare characteristics of temporal super-resolution to those of spatial super-resolution. These include: How many video cameras are needed to obtain increased resolution? What is the upper bound on resolution improvement via super-resolution? What is the temporal analogue to the spatial "ringing" effect?
Eli Shechtman, Yaron Caspi, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Space-Time Video Completion
Yonatan Wexler, Eli Shechtman, Michal Irani
CVPR (1)3
2004 Separating Transparent Layers through Layer Information Exchange
Bernard Sarel, Michal Irani
ECCV (4)2
2004 Temporal Factorization vs. Spatial Factorization
Lihi Zelnik-Manor, Michal Irani
ECCV (2)2
2003 Degeneracies, Dependencies and their Implications in Multi-body and Multi-Sequence Factorizations
abstract
The body of work on multi-body factorization separates between objects whose motions are independent. In this work we show that in many cases objects moving with different 3D motions will be captured as a single object using these approaches. We analyze what causes these degeneracies between objects and suggest an approach for overcoming some of them. We further show that in the case of multiple sequences linear dependencies can supply information for temporal synchronization of sequences and for spatial matching of points across sequences.
Lihi Zelnik-Manor, Michal Irani
CVPR (2)2
2002 What Does the Scene Look Like from a Scene Point?
Michal Irani, Tal Hassner, P. Anandan 0001
ECCV (2)1
2002 Increasing Space-Time Resolution in Video
Eli Shechtman, Yaron Caspi, Michal Irani
ECCV (1)3
2002 Factorization with Uncertainty
P. Anandan 0001, Michal Irani
Int. J. Comput. Vis.2
2002 Aligning Non-Overlapping Sequences
Yaron Caspi, Michal Irani
Int. J. Comput. Vis.2
2002 Multi-Frame Correspondence Estimation Using Subspace Constraints
Michal Irani
Int. J. Comput. Vis.1
2002 Spatio-Temporal Alignment of Sequences
abstract
This paper studies the problem of sequence-to-sequence alignment, namely, establishing correspondences in time and in space between two different video sequences of the same dynamic scene. The sequences are recorded by uncalibrated video cameras which are either stationary or jointly moving, with fixed (but unknown) internal parameters and relative intercamera external parameters. Temporal variations between image frames (such as moving objects or changes in scene illumination) are powerful cues for alignment, which cannot be exploited by standard image-to-image alignment techniques. We show that, by folding spatial and temporal cues into a single alignment framework, situations which are inherently ambiguous for traditional image-to-image alignment methods, are often uniquely resolved by sequence-to-sequence alignment. Furthermore, the ability to align and integrate information across multiple video sequences both in time and in space gives rise to new video applications that are not possible when only image-to-image alignment is used.
Yaron Caspi, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Direct Recovery of Planar-Parallax from Multiple Frames
abstract
We present an algorithm that estimates dense planar-parallax motion from multiple uncalibrated views of a 3D scene. This generalizes the "plane+parallax" recovery methods to more than two frames. The parallax motion of pixels across multiple frames (relative to a planar surface) is related to the 3D scene structure and the camera epipoles. The parallax field, the epipoles, and the 3D scene structure are estimated directly from image brightness variations across multiple frames, without precomputing correspondences.
Michal Irani, P. Anandan 0001, Meir Cohen
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Multiview Constraints on Homographies
abstract
The image motion of a planar surface between two camera views is captured by a homography (a 2D projective transformation). The homography depends on the intrinsic and extrinsic camera parameters, as well as on the 3D plane parameters. While camera parameters vary across different views, the plane geometry remains the same. Based on this fact, we derive linear subspace constraints on the relative homographies of multiple (/spl ges/ 2) planes across multiple views. The paper has three main contributions: 1) We show that the collection of all relative homographies (homologies) of a pair of planes across multiple views, spans a 4-dimensional linear subspace. 2) We show how this constraint can be extended to the case of multiple planes across multiple views. 3) We show that, for some restricted cases of camera motion, linear subspace constraints apply also to the set of homographies of a single plane across multiple views. All the results derived are true for uncalibrated cameras. The possible utility of these multiview constraints for improving homography estimation and for detecting nonrigid motions are also discussed.
Lihi Zelnik-Manor, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Event-Based Analysis of Video
abstract
Dynamic events can be regarded as long-term temporal objects, which are characterized by spatio-temporal features at multiple temporal scales. Based on this, we design a simple statistical distance measure between video sequences (possibly of different lengths) based on their behavioral content. This measure is non-parametric and can thus handle a wide range of dynamic events. We use this measure for isolating and clustering events within long continuous video sequences. This is done without prior knowledge of the types of events, their models, or their temporal extent. An outcome of such a clustering process is a temporal segmentation of long video sequences into event-consistent sub-sequences, and their grouping into event-consistent clusters. Our event representation and associated distance measure can also be used for event-based indexing into long video sequences, even when only one short example-clip is available. However, when multiple example-clips of the same event are available (either as a result of the clustering process, or given manually), these can be used to refine the event representation, the associated distance measure, and accordingly the quality of the detection and clustering process.
Lihi Zelnik-Manor, Michal Irani
CVPR (2)2
2001 Alignment of Non-Overlapping Sequences
abstract
This paper shows how two image sequences that have no spatial overlap between their fields of view can be aligned both in time and in space. Such alignment is possible when the two cameras are attached closely together and are moved jointly in space. The common motion induces "similar" changes over time within the two sequences. This correlated temporal behavior is used to recover the spatial and temporal transformations between the two sequences. The requirement of "coherent appearance" in standard image alignment techniques is therefore replaced by "coherent temporal behavior", which is often easier to satisfy. This approach to alignment can be used not only for aligning nan-overlapping sequences, but also for handling other cases that are inherently difficult for standard image alignment techniques. We demonstrate applications of this approach to three real-world problems: (i) alignment of non-overlapping sequences for generating wide-screen movies, (ii) alignment of images (sequences) obtained at significantly different zooms, for surveillance applications, and (iii) multi-sensor image alignment for multi-sensor fusion.
Yaron Caspi, Michal Irani
ICCV2
2000 Step towards Sequence-to-Sequence Alignment
abstract
The paper presents an approach for establishing correspondences in time and in space between two different video sequences of the same dynamic scene, recorded by stationary uncalibrated video cameras. The method simultaneously estimates both spatial alignment as well as temporal synchronization (temporal alignment) between the two sequences, using all available spatio-temporal information. Temporal variations between image frames (such as moving objects or changes in scene illumination) are powerful cues for alignment, which cannot be exploited by standard image-to-image alignment techniques. We show that by folding spatial and temporal cues into a single alignment framework, situations which are inherently ambiguous for traditional image-to-image alignment methods, are often uniquely resolved by sequence-to-sequence alignment. We also present a "direct" method for sequence-to-sequence alignment. The algorithm simultaneously estimates spatial and temporal alignment parameters directly from measurable sequence quantities, without requiring prior estimation of point correspondences, frame correspondences, or moving object detection. Results are shown on real image sequences taken by multiple video cameras.
Yaron Caspi, Michal Irani
CVPR2
2000 Factorization with Uncertainty
Michal Irani, P. Anandan 0001
ECCV (1)1
2000 Multi-Frame Estimation of Planar Motion
abstract
Traditional plane alignment techniques are typically performed between pairs of frames. We present a method for extending existing two-frame planar motion estimation techniques into a simultaneous multi-frame estimation, by exploiting multi-frame subspace constraints of planar surfaces. The paper has three main contributions: 1) we show that when the camera calibration does not change, the collection of all parametric image motions of a planar surface in the scene across multiple frames is embedded in a low dimensional linear subspace; 2) we show that the relative image motion of multiple planar surfaces across multiple frames is embedded in a yet lower dimensional linear subspace, even with varying camera calibration; and 3) we show how these multi-frame constraints can be incorporated into simultaneous multi-frame estimation of planar motion, without explicitly recovering any 3D information, or camera calibration. The resulting multi-frame estimation process is more constrained than the individual two-frame estimations, leading to more accurate alignment, even when applied to small image regions.
Lihi Zelnik-Manor, Michal Irani
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Multi-Frame Alignment of Planes
abstract
Traditional plane alignment techniques are typically performed between pairs of frames. In this paper we present a method for extending existing two-frame planar-motion estimation techniques into a simultaneous multi-frame estimation, by exploiting multi-frame geometric constraints of planar surfaces. The paper has three main contributions: (i) we show that when the camera calibration does not change, the collection of all parametric image motions of a planar surface in the scene across multiple frames is embedded in a low dimensional linear subspace; (ii) we show that the relative image motion of multiple planar surfaces across multiple frames is embedded in a yet lower dimensional linear subspace, even with varying camera calibration; and (iii) we show how these multi-frame constraints can be incorporated into simultaneous multi-frame estimation of planar motion, without explicitly recovering any 3D information, or camera calibration. The resulting multi-frame estimation process is more constrained than the individual two-frame estimations, leading to more accurate alignment, even when applied to small image regions.
Lihi Zelnik-Manor, Michal Irani
CVPR2
1999 Multi-Frame Optical Flow Estimation using Subspace Constraints
abstract
Shows that the set of all flow fields in a sequence of frames imaging a rigid scene resides in a low-dimensional linear subspace. Based on this observation, we develop a method for simultaneous estimation of optical flow across multiple frames, which uses these subspace constraints. The multi-frame subspace constraints are strong constraints, and they replace commonly used heuristic constraints, such as spatial or temporal smoothness. The subspace constraints are geometrically meaningful and are not violated at depth discontinuities or when the camera motion changes abruptly. Furthermore, we show that the subspace constraints on flow fields apply for a variety of imaging models, scene models and motion models. Hence, the presented approach for constrained multi-frame flow estimation is general. However, our approach does not require prior knowledge of the underlying world or camera model. Although linear subspace constraints have been used successfully in the past for recovering 3D information, it has been assumed that 2D correspondences are given. However, correspondence estimation is a fundamental problem in motion analysis. In this paper, we use multi-frame subspace constraints to constrain the 2D correspondence estimation process itself, and not for 3D recovery.
Michal Irani
ICCV1
1999 Multi-View Subspace Constraints on Homographies
abstract
The motion of a planar surface between two camera views induces a homography. The homography depends on the camera intrinsic and extrinsic parameters, as well as on the 3D plane parameters. While camera parameters vary across different views, the plane geometry remains the same. Based on this fact, the paper derives linear subspace constraints on the relative motion of multiple (/spl ges/2) planes across multiple views. The paper has three main contributions. It shows that the collection of all relative homographies of a pair of planes (homologies) across multiple views, spans a 4-dimensional linear subspace. It shows how this constraint can be extended to the case of multiple planes across multiple views. It suggests two potential application areas which can benefit from these constraints: the accuracy of homography estimation can be improved by enforcing the multi-view subspace constraints; and violations of these multi-view constraints can be used as a cue for moving object detection. All the results derived in this paper are true for uncalibrated cameras.
Lihi Zelnik-Manor, Michal Irani
ICCV2
1998 From Reference Frames to Reference Planes: Multi-View Parallax Geometry and Applications
Michal Irani, P. Anandan 0001, Daphna Weinshall
ECCV (2)1
1998 Robust Multi-Sensor Image Alignment
abstract
This paper presents a method for alignment of images acquired by sensors of different modalities (e.g., EO and IR). The paper has two main contributions: (i) It identifies an appropriate image representation, for multi-sensor alignment, i.e., a representation which emphasizes the common information between the two multi-sensor images, suppresses the non-common information, and is adequate for coarse-to-fine processing. (ii) It presents a new alignment technique which applies global estimation to any choice of a local similarity measure. In particular, it is shown that when this registration technique is applied to the chosen image representation with a local normalized-correlation similarity measure, it provides a new multi-sensor alignment algorithm which is robust to outliers, and applies to a wide variety of globally complex brightness transformations between the two images. Our proposed image representation does not rely on sparse image features (e.g., edge, contour, or point features). It is continuous and does not eliminate the detailed variations within local image regions. Our method naturally extends to coarse-to-fine processing, and applies even in situations when the multi-sensor signals are globally characterized by low statistical correlation.
Michal Irani, P. Anandan 0001
ICCV1
1998 A Unified Approach to Moving Object Detection in 2D and 3D Scenes
abstract
The detection of moving objects is important in many tasks. Previous approaches to this problem can be broadly divided into two classes: 2D algorithms which apply when the scene can be approximated by a flat surface and/or when the camera is only undergoing rotations and zooms, and 3D algorithms which work well only when significant depth variations are present in the scene and the camera is translating. We describe a unified approach to handling moving object detection in both 2D and 3D scenes, with a strategy to gracefully bridge the gap between those two extremes. Our approach is based on a stratification of the moving object detection problem into scenarios which gradually increase in their complexity. We present a set of techniques that match the above stratification. These techniques progressively increase in their complexity, ranging from 2D techniques to more complex 3D techniques. Moreover, the computations required for the solution to the problem at one complexity level become the initial processing step for the solution at the next complexity level. We illustrate these techniques using examples from real-image sequences.
Michal Irani, P. Anandan 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Video indexing based on mosaic representations
abstract
Video is a rich source of information. It provides visual information about scenes. This information is implicitly buried inside the raw video data, however, and is provided with the cost of very high temporal redundancy. While the standard sequential form of video storage is adequate for viewing in a movie mode, it fails to support rapid access to information of interest that is required in many of the emerging applications of video. This paper presents an approach for efficient access, use and manipulation of video data. The video data are first transformed from their sequential and redundant frame-based representation, in which the information about the scene is distributed over many frames, to an explicit and compact scene-based representation, to which each frame can be directly related. This compact reorganization of the video data supports nonlinear browsing and efficient indexing to provide rapid access directly to information of interest. This paper describes a new set of methods for indexing into the video sequence based on the scene-based representation. These indexing methods are based on geometric and dynamic information contained in the video. These methods complement the more traditional content-based indexing methods, which utilize image appearance information (namely, color and texture properties) but are considerably simpler to achieve and are highly computationally efficient.
Michal Irani, P. Anandan 0001
Proc. IEEE1
1997 Interactive content-based video indexing and browsing
abstract
In this paper we present a framework for efficient representation, access, and manipulation of video data. Our approach is based on decomposing video information into its spatial (appearance), temporal (dynamics), geometric components. This derived information is organized into data representations that support non-linear browsing and efficient indexing to provide rapid access directly to the information of interest.
Michal Irani, Harpreet Sawhney, Rakesh Kumar 0001, P. Anandan 0001
MMSP1
1997 Recovery of Ego-Motion Using Region Alignment
abstract
A method for computing the 3D camera motion (the ego-motion) in a static scene is described, where initially a detected 2D motion between two frames is used to align corresponding image regions. We prove that such a 2D registration removes all effects of camera rotation, even for those image regions that remain misaligned. The resulting residual parallax displacement field between the two region-aligned images is an epipolar field centered at the FOE (Focus-of-Expansion). The 3D camera translation is recovered from the epipolar field. The 3D camera rotation is recovered from the computed 3D translation and the detected 2D motion. The decomposition of image motion into a 2D parametric motion and residual epipolar parallax displacements avoids many of the inherent ambiguities and instabilities associated with decomposing the image motion into its rotational and translational components, and hence makes the computation of ego-motion or 3D structure estimation more robust.
Michal Irani, Benny Rousso, Shmuel Peleg
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Parallax Geometry of Pairs of Points for 3D Scene Analysis
Michal Irani, P. Anandan 0001
ECCV (1)1
1996 A unified approach to moving object detection in 2D and 3D scenes
abstract
The detection of moving objects is important in many tasks. Previous approaches to this problem can be broadly divided into two classes: 2D algorithms which apply when the scene can be approximated by a flat surface and/or when the camera is only undergoing rotations and zooms; and 3D algorithms which work well only when significant depth variations are present in the scene and the camera is translating. In this paper, we describe a unified approach to handling moving object detection in both 2D and 3D scenes, with a strategy to gracefully bridge the gap between those two extremes. Our approach is based on a stratification of the moving object detection problem into scenarios and corresponding techniques which gradually increase in their complexity. Moreover, the computations required for the solution to the problem at one complexity level become the initial processing step for the solution at the next complexity level.
Michal Irani, P. Anandan 0001
ICPR1
1996 Efficient representations of video sequences and their applications
Michal Irani, P. Anandan 0001, James R. Bergen, Rakesh Kumar 0001, Steven C. Hsu
Signal Process. Image Commun.1
1995 Mosaic Based Representations of Video Sequences and Their Applications
abstract
Recently, there has been a growing interest in the use of mosaic images to represent the information contained in video sequences. The paper systematically investigates how to go beyond thinking of the mosaic simply as a visualization device, but rather as a basis for efficient representation of video sequences. We describe two different types of mosaics called the static and the dynamic mosaic that are suitable for different needs and scenarios. We discuss a series of extensions to these basic mosaics to provide representations at multiple spatial and temporal resolutions and to handle 3D scene information. We describe techniques for the basic elements of the mosaic construction process, namely alignment, integration, and residual analysis. We describe several applications of mosaic representations including video compression, enhancement, enhanced visualization, and other applications in video indexing, search, and manipulation.>
Michal Irani, P. Anandan 0001, Steven C. Hsu
ICCV1
1995 Video as an image data source: efficient representations and applications
abstract
The two fundamental advantages of video over still imagery are: (i) the ability capture temporal information, and (ii) the ability to acquire a continuously varying set of views of a scene. These advantages are obtained, however, at the cost of vastly increased amount of data. This paper describes an approach to video representation that is based on frame-to-frame alignment, mosaic construction, and 3D parallax recovery. The basic motivation behind our approach is to enable rapid access to the contents, while maintaining the data in a form as close to the source as possible. This representation supports a wide variety of applications that involve transmission, storage, visualization, retrieval, analysis, and manipulation of video sequences.
P. Anandan 0001, Michal Irani, Rakesh Kumar 0001, James R. Bergen
ICIP2
1995 Video compression using mosaic representations
Michal Irani, Steven C. Hsu, P. Anandan 0001
Signal Process. Image Commun.1
1994 Recovery of ego-motion using image stabilization
abstract
A method for computing the 3D camera motion (the ego-motion) in a static scene is introduced, which is based on computing the 2D image motion of a single image region directly from image intensities. The computed image motion of this image region is used to register the images so that the detected image region appears stationary. The resulting displacement field for the entire scene between the registered frames is affected only by the 3D translation of the camera. After canceling the effects of the camera rotation by using such 2D image registration, the 3D camera translation is computed by finding the focus-of-expansion in the translation-only set of registered frames. This step is followed by computing the camera rotation to complete the computation of the ego-motion. The presented method avoids the inherent problems in the computation of optical flow and of feature matching, and does not assume any prior feature detection or feature correspondence.>
Michal Irani, Benny Rousso, Shmuel Peleg
CVPR1
1994 Computing occluding and transparent motions
Michal Irani, Benny Rousso, Shmuel Peleg
Int. J. Comput. Vis.1
1993 Robust Recovery of Ego-Motion
Michal Irani, Benny Rousso, Shmuel Peleg
CAIP1
1993 Motion Analysis for Image Enhancement: Resolution, Occlusion, and Transparency
Michal Irani, Shmuel Peleg
J. Vis. Commun. Image Represent.1
1992 Image sequence enhancement using multiple motions analysis
abstract
A method for detecting and tracking multiple moving objects, using both a large spatial region and a large temporal region, without assuming temporal motion constancy is described. When the large spatial region of analysis has multiple moving objects, the motion parameters and the locations of the objects are computed for one object after another. A method for segmenting the image plane into differently moving objects and computing their motions using two frames is presented. The tracking of detected objects using temporal integration and the algorithms for enhancement of tracked objects by filling-in occluded regions and by improving the spatial resolution of the imaged objects are described.>
Michal Irani, Shmuel Peleg
CVPR1
1992 Detecting and Tracking Multiple Moving Objects Using Temporal Integration
Michal Irani, Benny Rousso, Shmuel Peleg
ECCV1
1991 Improving resolution by image registration
Michal Irani, Shmuel Peleg
CVGIP Graph. Model. Image Process.1
1990 Super resolution from image sequences
abstract
An iterative algorithm to increase image resolution is described. Examples are shown for low-resolution gray-level pictures, with an increase of resolution clearly observed after only a few iterations. The same method can also be used for deblurring a single blurred image. The approach is based on the resemblance of the presented problem to the reconstruction of a 2-D object from its 1-D projections in computer-aided tomography. The algorithm performed well for both computer-simulated and real images and is shown, theoretically and practically, to converge quickly. The algorithm can be executed in parallel for faster hardware implementation.>
Michal Irani, Shmuel Peleg
ICPR (2)1