Yosi Keller

dblp:41/4110 · DBLP profile ↗
← Back
53ranked-venue papers
14as first author
18since 2021 · last 2026
0000-0002-2876-2790ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 first-author · 6 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Beyond familiar landscapes: Exploring the limits of relative pose regressors in new environments
abstract
Relative pose regressors (RPRs) determine the pose of a query image by estimating its relative translation and rotation to a reference pose-labeled camera. Unlike other regression-based localization techniques confined to a scene’s absolute parameters, RPRs learn residuals, making them adaptable to new environments. However, RPRs have exhibited limited generalization to scenes not utilized during training (“unseen scenes”). In this work, we explore the ability of RPRs to localize in unseen scenes and propose algorithmic modifications to enhance their generalization. These modifications include attention-based aggregation of coarse feature maps, dynamic adaptation of model weights, and geometry-aware optimization. Our proposed approach improves the localization accuracy of RPRs in unseen scenes by a notable margin across multiple indoor and outdoor benchmarks and under various conditions while maintaining comparable performance in scenes used during training. We assess the contribution of each component through ablation studies and further analyze the uncertainty of our model in unseen scenes. Our Code and pre-trained models are available at https://github.com/yolish/relformer .
Ofer Idan, Yoli Shavit, Yosi Keller
Comput. Vis. Image Underst.3
2025 HyperPose: Hypernetwork-Infused Camera Pose Localization and an Extended Cambridge Landmarks Dataset
abstract
In this work, we propose HyperPose, which utilizes hypernetworks in absolute camera pose regressors. The inherent appearance variations in natural scenes, attributable to environmental conditions, perspective, and lighting, induce a significant domain disparity between the training and test datasets. This disparity degrades the precision of contemporary localization networks. To mitigate this, we advocate for incorporating hypernetworks into single-scene and multiscene camera pose regression models. During inference, the hypernetwork dynamically computes adaptive weights for the localization regression heads based on the particular input image, effectively narrowing the domain gap. Using indoor and outdoor datasets, we evaluate the HyperPose methodology across multiple established absolute pose regression architectures. We also introduce and share the Extended Cambridge Landmarks (ECL), a novel localization dataset, based on the Cambridge Landmarks dataset, showing it in multiple seasons with significantly varying appearance conditions. Our empirical experiments demonstrate that HyperPose yields notable performance enhancements for single- and multi-scene architectures. We have made our source code, pre-trained models1, and the ECL dataset openly available2.
Ron Ferens, Yosi Keller
CVPR2
2024 FaceCoresetNet: Differentiable Coresets for Face Set Recognition
abstract
In set-based face recognition, we aim to compute the most discriminative descriptor from an unbounded set of images and videos showing a single person. A discriminative descriptor balances two policies when aggregating information from a given set. The first is a quality-based policy: emphasizing high-quality and down-weighting low-quality images. The second is a diversity-based policy: emphasizing unique images in the set and down-weighting multiple occurrences of similar images as found in video clips which can overwhelm the set representation. This work frames face-set representation as a differentiable coreset selection problem. Our model learns how to select a small coreset of the input set that balances quality and diversity policies using a learned metric parameterized by the face quality, optimized end-to-end. The selection process is a differentiable farthest-point sampling (FPS) realized by approximating the non-differentiable Argmax operation with differentiable sampling from the Gumbel-Softmax distribution of distances. The small coreset is later used as queries in a self and cross-attention architecture to enrich the descriptor with information from the whole set. Our model is order-invariant and linear in the input set size. We set a new SOTA to set face verification on the IJB-B and IJB-C datasets. Our code is publicly available at https://github.com/ligaripash/FaceCoresetNet.
Gil Shapira, Yosi Keller
AAAI2
2024 Estimating Extreme 3D Image Rotations using Cascaded Attention
abstract
Estimating large, extreme inter-image rotations is crit-ical for numerous computer vision domains involving images related by limited or non-overlapping fields of view. In this work, we propose an attention-based approach with a pipeline of novel algorithmic components. First, as ro-tation estimation pertains to image pairs, we introduce an inter-image distillation scheme using Decoders to improve embeddings. Second, whereas contemporary methods com-pute a 4D correlation volume (4DCV) encoding inter-image relationships, we propose an Encoder-based cross-attention approach between activation maps to compute an enhanced equivalent of the 4DCV. Finally, we present a cascaded Decoder-based technique for alternately refining the cross-attention and the rotation query. Our approach outperforms current state-of-the-art methods on extreme rotation estimation. We make our code publicly available11https://github.com/dekelshay/AttExtremeRotation.
Shay Dekel, Yosi Keller, Martin Cadík
CVPR2
2024 Attention-based multimodal image matching
Aviad Moreshet, Yosi Keller
Comput. Vis. Image Underst.2
2024 Learning single and multi-scene camera pose regression with transformer encoders
Yoli Shavit, Ron Ferens, Yosi Keller
Comput. Vis. Image Underst.3
2024 Deep Convolutional Tables: Deep Learning Without Convolutions
abstract
We propose a novel formulation of deep networks that do not use dot-product neurons and rely on a hierarchy of voting tables instead, denoted as convolutional tables (CTs), to enable accelerated CPU-based inference. Convolutional layers are the most time-consuming bottleneck in contemporary deep learning techniques, severely limiting their use in the Internet of Things and CPU-based devices. The proposed CT performs a fern operation at each image location: it encodes the location environment into a binary index and uses the index to retrieve the desired local output from a table. The results of multiple tables are combined to derive the final output. The computational complexity of a CT transformation is independent of the patch (filter) size and grows gracefully with the number of channels, outperforming comparable convolutional layers. It is shown to have a better capacity:compute ratio than dot-product neurons, and that deep CT networks exhibit a universal approximation property similar to neural networks. As the transformation involves computing discrete indices, we derive a soft relaxation and gradient-based approach for training the CT hierarchy. Deep CT networks have been experimentally shown to have accuracy comparable to that of CNNs of similar architectures. In the low-compute regime, they enable an error:speed tradeoff superior to alternative efficient CNN architectures.
Shay Dekel, Yosi Keller, Aharon Bar-Hillel
IEEE Trans. Neural Networks Learn. Syst.2
2023 Vision UFormer: Long-range monocular absolute depth estimation
Tomas Polasek, Martin Cadík, Yosi Keller, Bedrich Benes
Comput. Graph.3
2023 Hierarchical Attention-Based Age Estimation and Bias Analysis
abstract
In this work, we present a Deep Learning approach to estimate age from facial images. First, we introduce a novel attention-based approach to image augmentation-aggregation, which allows multiple image augmentations to be adaptively aggregated using a Transformer-Encoder. A hierarchical probabilistic regression model is then proposed that combines discrete probabilistic age estimates with an ensemble of regressors. Each regressor is adapted and trained to refine the probability estimate over a given age range. We show that our age estimation scheme outperforms current schemes and provides a new state-of-the-art age estimation accuracy when applied to the MORPH II and CACD datasets. We also present an analysis of the biases in the results of the state-of-the-art age estimates.
Shakediel Hiba, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 FSGANv2: Improved Subject Agnostic Face Swapping and Reenactment
abstract
We present Face Swapping GAN (FSGAN) for face swapping and reenactment. Unlike previous work, we offer a subject agnostic swapping scheme that can be applied to pairs of faces without requiring training on those faces. We derive a novel iterative deep learning-based approach for face reenactment which adjusts significant pose and expression variations that can be applied to a single image or a video sequence. For video sequences, we introduce a continuous interpolation of the face views based on reenactment, Delaunay Triangulation, and barycentric coordinates. Occluded face regions are handled by a face completion network. Finally, we use a face blending network for seamless blending of the two faces while preserving the target skin color and lighting conditions. This network uses a novel Poisson blending loss combining Poisson optimization with a perceptual loss. We compare our approach to existing state-of-the-art systems and show our results to be both qualitatively and quantitatively superior. This work describes extensions of the FSGAN method, proposed in an earlier conference version of our work (Nirkin et al. 2019), as well as additional experiments and results.
Yuval Nirkin, Yosi Keller, Tal Hassner
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Coarse-to-Fine Multi-Scene Pose Regression With Transformers
abstract
Absolute camera pose regressors estimate the position and orientation of a camera given the captured image alone. Typically, a convolutional backbone with a multi-layer perceptron (MLP) head is trained using images and pose labels to embed a single reference scene at a time. Recently, this scheme was extended to learn multiple scenes by replacing the MLP head with a set of fully connected layers. In this work, we propose to learn multi-scene absolute camera pose regression with Transformers, where encoders are used to aggregate activation maps with self-attention and decoders transform latent features and scenes encoding into pose predictions. This allows our model to focus on general features that are informative for localization, while embedding multiple scenes in parallel. We extend our previous MS-Transformer approach Shavit et al. (2021) by introducing a mixed classification-regression architecture that improves the localization accuracy. Our method is evaluated on commonly benchmark indoor and outdoor datasets and has been shown to exceed both multi-scene and state-of-the-art single-scene absolute pose regressors.
Yoli Shavit, Ron Ferens, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Camera Pose Auto-encoders for Improving Pose Regression
Yoli Shavit, Yosi Keller
ECCV (10)2
2022 Joint Detection and Matching of Feature Points in Multimodal Images
abstract
In this work, we propose a novel Convolutional Neural Network (CNN) architecture for the joint detection and matching of feature points in images acquired by different sensors using a single forward pass. The resulting feature detector is tightly coupled with the feature descriptor, in contrast to classical approaches (SIFT, etc.), where the detection phase precedes and differs from computing the descriptor. Our approach utilizes two CNN subnetworks, the first being a Siamese CNN and the second, consisting of dual non-weight-sharing CNNs. This allows simultaneous processing and fusion of the joint and disjoint cues in the multimodal image patches. The proposed approach is experimentally shown to outperform contemporary state-of-the-art schemes when applied to multiple datasets of multimodal images. It is also shown to provide repeatable feature points detections across multi-sensor images, outperforming state-of-the-art detectors. To the best of our knowledge, it is the first unified approach for the detection and matching of such images.
Elad Ben Baruch, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Learning to Embed Semantic Similarity for Joint Image-Text Retrieval
abstract
We present a deep learning approach for learning the joint semantic embeddings of images and captions in a euclidean space, such that the semantic similarity is approximated by the$L_{2}$distances in the embedding space. For that, we introduce a metric learning scheme that utilizes multitask learning to learn the embedding ofidenticalsemantic concepts using a center loss. By introducing a differentiable quantization scheme into the end-to-end trainable network, we derive a semantic embedding of semanticallysimilarconcepts in euclidean space. We also propose a novel metric learning formulation using an adaptive margin hinge loss, that is refined during the training phase. The proposed scheme was applied to the MS-COCO, Flicke30K and Flickr8K datasets, and was shown to compare favorably with contemporary state-of-the-art approaches.
Noam Malali, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 DeepFake Detection Based on Discrepancies Between Faces and Their Context
abstract
We propose a method for detecting face swapping and other identity manipulations in single images. Face swapping methods, such as DeepFake, manipulate the face region, aiming to adjust the face to the appearance of its context, while leaving the context unchanged. We show that this modus operandi produces discrepancies between the two regions (e.g., Fig. 1). These discrepancies offer exploitable telltale signs of manipulation. Our approach involves two networks: (i) a face identification network that considers the face region bounded by a tight semantic segmentation, and (ii) a context recognition network that considers the face context (e.g., hair, ears, neck). We describe a method which uses the recognition signals from our two networks to detect such discrepancies, providing a complementary detection signal that improves conventional real versus fake classifiers commonly used for detecting fake images. Our method achieves state of the art results on the FaceForensics++ and Celeb-DF-v2 benchmarks for face manipulation detection, and even generalizes to detect fakes produced by unseen methods.
Yuval Nirkin, Lior Wolf, Yosi Keller, Tal Hassner
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Learning Multi-Scene Absolute Pose Regression with Transformers
abstract
Absolute camera pose regressors estimate the position and orientation of a camera from the captured image alone. Typically, a convolutional backbone with a multi-layer perceptron head is trained using images and pose labels to embed a single reference scene at a time. Recently, this scheme was extended for learning multiple scenes by replacing the MLP head with a set of fully connected layers. In this work, we propose to learn multi-scene absolute camera pose regression with Transformers, where encoders are used to aggregate activation maps with self-attention and decoders transform latent features and scenes encoding into candidate pose predictions. This mechanism allows our model to focus on general features that are informative for localization while embedding multiple scenes in parallel. We evaluate our method on commonly benchmarked indoor and outdoor datasets and show that it surpasses both multi-scene and state-of-the-art single-scene absolute pose regressors. We make our code publicly available from https://github.com/yolish/multi-scene-pose-transformer.
Yoli Shavit, Ron Ferens, Yosi Keller
ICCV3
2021 Facial landmarks localization using cascaded neural networks
Shahar Mahpod, Rig Das, Emanuele Maiorana, Yosi Keller, Patrizio Campisi
Comput. Vis. Image Underst.4
2021 A Unified Approach to Kinship Verification
abstract
In this work, we propose a deep learning-based approach for kin verification using a unified multi-task learning scheme where all kinship classes are jointly learned. This allows us to better utilize small training sets that are typical of kin verification. We introduce a novel approach for fusing the embeddings of kin images, to avoid overfitting, which is a common issue in training such networks. An adaptive sampling scheme is derived for the training set images, to resolve the inherent imbalance in kin verification datasets. A thorough ablation study exemplifies the effectivity of our approach, which is experimentally shown to outperform contemporary state-of-the-art kin verification results when applied to the Families In the Wild, FG2018, and FG2020 datasets.
Eran Dahan, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Multi-scale Processing of Noisy Images using Edge Preservation Losses
abstract
Noisy image processing is a fundamental task of computer vision. The first example is the detection of faint edges in noisy images, a challenging problem studied in the last decades. A recent study introduced a fast method to detect faint edges in the highest accuracy among all the existing approaches. Their complexity is nearly linear in the image's pixels and their runtime is seconds for a noisy image. Their approach utilizes a multiscale binary partitioning of the image. By utilizing the multiscale U-net architecture, we show in this paper that their method can be dramatically improved in both aspects of run time and accuracy. By training the network on a dataset of binary images, we developed an approach for faint edge detection that works in linear complexity. Our runtime of a noisy image is milliseconds on a GPU. Even though our method is orders of magnitude faster, we still achieve higher accuracy of detection under many challenging scenarios. In addition, we show that our approach to performing multi-scale preprocessing of noisy images using U-net improves the ability to perform other vision tasks under the presence of noise. We prove it on the problems of noisy objects classification and classical image denoising. We show that multi-scale denoising can be carried out by a novel edge preservation loss. As our experiments show, we achieve high-quality results in the three aspects of faint edge detection, noisy image classification, and natural image denoising.
Nati Ofir, Yosi Keller
ICPR2
2019 FSGAN: Subject Agnostic Face Swapping and Reenactment
abstract
We present Face Swapping GAN (FSGAN) for face swapping and reenactment. Unlike previous work, FSGAN is subject agnostic and can be applied to pairs of faces without requiring training on those faces. To this end, we describe a number of technical contributions. We derive a novel recurrent neural network (RNN)-based approach for face reenactment which adjusts for both pose and expression variations and can be applied to a single image or a video sequence. For video sequences, we introduce continuous interpolation of the face views based on reenactment, Delaunay Triangulation, and barycentric coordinates. Occluded face regions are handled by a face completion network. Finally, we use a face blending network for seamless blending of the two faces while preserving target skin color and lighting conditions. This network uses a novel Poisson blending loss which combines Poisson optimization with perceptual loss. We compare our approach to existing state-of-the-art systems and show our results to be both qualitatively and quantitatively superior.
Yuval Nirkin, Yosi Keller, Tal Hassner
ICCV2
2019 Iterative spectral independent component analysis
Shai Gepshtein, Yosi Keller
Signal Process.2
2019 DeepAge: Deep Learning of face-based age estimation
Omry Sendik, Yosi Keller
Signal Process. Image Commun.2
2018 Deep Multi-Spectral Registration Using Invariant Descriptor Learning
abstract
In this work, we propose a deep-learning approach for aligning cross-spectral images. Our approach utilizes a learned descriptor invariant to different spectra. Multi-modal images of the same scene capture different characteristics and therefore their registration is challenging. To that end, we developed a feature-based approach for registering visible (VIS) to Near-Infra-Red (NIR) images. Our scheme detects corners by Harris and matches them by a patch-metric learned on top of a network trained using the CIFAR-10 dataset. As our experiments demonstrate, we achieve accurate alignment of cross-spectral images with sub-pixel accuracy. Comparing to contemporary state-of-the-art, our approach is more accurate in the task of VIS to NIR registration.
Nati Ofir, Shai Silberstein, Hila Levi, Dani Rozenbaum, Yosi Keller, Sharon Duvdevani Bar
ICIP5
2018 Registration and Fusion of Multi-Spectral Images Using a Novel Edge Descriptor
abstract
In this work we propose a fully end-to-end approach for multi-spectral image registration and fusion. Our fusion method combines images from different spectral channels into a single fused image using approaches for low and high frequency signals. A prerequisite of fusion is the geometric alignment between the spectral bands, commonly referred to as registration. Unfortunately, common methods for image registration of a single spectral channel might prove inaccurate on images from different modalities. For that end, we introduce a new algorithm for multi-spectral image registration, based on a novel edge descriptor of feature points. Our method achieves an accurate alignment allowing us to further fuse the images. It is experimentally shown to produce a high quality of multi-spectral image registration and fusion under challenging scenarios.
Nati Ofir, Shai Silberstein, Dani Rozenbaum, Yosi Keller, Sharon Duvdevani Bar
ICIP4
2018 Kinship verification using multiview hybrid distance learning
Shahar Mahpod, Yosi Keller
Comput. Vis. Image Underst.2
2016 Image Segmentation via Probabilistic Graph Matching
abstract
This paper presents an unsupervised and semi-automatic image segmentation approach where we formulate the segmentation as an inference problem based on unary and pairwise assignment probabilities computed using low-level image cues. The inference is solved via a probabilistic graph matching scheme, which allows rigorous incorporation of low-level image cues and automatic tuning of parameters. The proposed scheme is experimentally shown to compare favorably with contemporary semi-supervised and unsupervised image segmentation schemes, when applied to contemporary state-of-the-art image sets.
Ayelet Heimowitz, Yosi Keller
IEEE Trans. Image Process.2
2016 An Algorithm for Improving Non-Local Means Operators via Low-Rank Approximation
abstract
We present a method for improving a non-local means (NLM) operator by computing its low-rank approximation. The low-rank operator is constructed by applying a filter to the spectrum of the original NLM operator. This results in an operator, which is less sensitive to noise while preserving important properties of the original operator. The method is efficiently implemented based on Chebyshev polynomials and is demonstrated on the application of natural images denoising. For this application, we provide a comparison of our method with other denoising methods.
Victor May, Yosi Keller, Nir Sharon, Yoel Shkolnisky
IEEE Trans. Image Process.2
2014 A Spectral Approach to Inter-Carrier Interference Mitigation in OFDM Systems
abstract
In this paper, we propose a new method for inter-carrier interference (ICI) mitigation in orthogonal frequency-division multiplexing (OFDM) systems. The proposed approach views the signal reconstruction problem at the receiver end as an integer least squares (ILS) problem, and uses a recently developed spectral approach called sequential probabilistic ILS (SPILS) to solve it. The proposed approach outperforms other state-of-the-art approaches while having the same computational complexity. In addition, we present a novel extension to the SPILS scheme that allows the generation of soft decisions (for communication systems which use soft-decision decoding). The use of soft-decision decoding (naturally) brings significant improvement in the detection reliability, and we show that the proposed method again outperforms other state-of-the-art approaches. To better address the tradeoff between performance and complexity, we first suggest a novel method to reduce the number of matrix inversions required and hence, to reduce the implementation complexity without any degradation in performance. We also introduce a novel low complexity scheme termed Quick SPILS (QSPILS) in which we lose a little in detection reliability, but significantly reduce the implementation complexity.
Avi Septimus, Yosi Keller, Itsik Bergel
IEEE Trans. Commun.2
2014 A Probabilistic Graph-Based Framework for Plug-and-Play Multi-Cue Visual Tracking
abstract
In this paper, we propose a novel approach for integrating multiple tracking cues within a unified probabilistic graph-based Markov random fields (MRFs) representation. We show how to integrate temporal and spatial cues encoded by unary and pairwise probabilistic potentials. As the inference of such high-order MRF models is known to be NP-hard, we propose an efficient spectral relaxation-based inference scheme. The proposed scheme is exemplified by applying it to a mixture of five tracking cues, and is shown to be applicable to wider sets of cues. This paves the way for a modular plug-and-play tracking framework that can be easily adapted to diverse tracking scenarios. The proposed scheme is experimentally shown to compare favorably with contemporary state-of-the-art schemes, and provides accurate tracking results.
Shimrit Feldman-Haber, Yosi Keller
IEEE Trans. Image Process.2
2013 A Probabilistic Approach to Spectral Graph Matching
abstract
Spectral Matching (SM) is a computationally efficient approach to approximate the solution of pairwise matching problems that are np-hard. In this paper, we present a probabilistic interpretation of spectral matching schemes and derive a novel Probabilistic Matching (PM) scheme that is shown to outperform previous approaches. We show that spectral matching can be interpreted as a Maximum Likelihood (ML) estimate of the assignment probabilities and that the Graduated Assignment (GA) algorithm can be cast as a Maximum a Posteriori (MAP) estimator. Based on this analysis, we derive a ranking scheme for spectral matchings based on their reliability, and propose a novel iterative probabilistic matching algorithm that relaxes some of the implicit assumptions used in prior works. We experimentally show our approaches to outperform previous schemes when applied to exhaustive synthetic tests as well as the analysis of real image sequences.
Amir Egozi, Yosi Keller, Hugo Guterman
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Image Completion by Diffusion Maps and Spectral Relaxation
abstract
We present a framework for image inpainting that utilizes the diffusion framework approach to spectral dimensionality reduction. We show that on formulating the inpainting problem in the embedding domain, the domain to be inpainted is smoother in general, particularly for the textured images. Thus, the textured images can be inpainted through simple exemplar-based and variational methods. We discuss the properties of the induced smoothness and relate it to the underlying assumptions used in contemporary inpainting schemes. As the diffusion embedding is nonlinear and noninvertible, we propose a novel computational approach to approximate the inverse mapping from the inpainted embedding space to the image domain. We formulate the mapping as a discrete optimization problem, solved through spectral relaxation. The effectiveness of the presented method is exemplified by inpainting real images, where it is shown to compare favorably with contemporary state-of-the-art schemes.
Shai Gepshtein, Yosi Keller
IEEE Trans. Image Process.2
2012 Scale-Invariant Features for 3-D Mesh Models
abstract
In this paper, we present a framework for detecting interest points in 3-D meshes and computing their corresponding descriptors. For that, we propose an intrinsic scale detection scheme per interest point and utilize it to derive two scale-invariant local features for mesh models. First, we present the scale-invariant spin image local descriptor that is a scale-invariant formulation of the spin image descriptor. Second, we adapt the scale-invariant feature transform feature to mesh data by representing the vicinity of each interest point as a depth map and estimating its dominant angle using the principal component analysis to achieve rotation invariance. The proposed features were experimentally shown to be robust to scale changes and partial mesh matching, and they were compared favorably with other local mesh features on the SHREC'10 and SHREC'11 testbeds. We applied the proposed local features to mesh retrieval using the bag-of-features approach and achieved state-of-the-art retrieval accuracy. Last, we applied the proposed local features to register models to scanned depth scenes and achieved high registration accuracy.
Tal Darom, Yosi Keller
IEEE Trans. Image Process.2
2010 3-D Symmetry Detection and Analysis Using the Pseudo-polar Fourier Transform
Amit Bermanis, Amir Averbuch, Yosi Keller
Int. J. Comput. Vis.3
2010 Spectral Symmetry Analysis
abstract
We present a spectral approach for detecting and analyzing rotational and reflectional symmetries in n-dimensions. Our main contribution is the derivation of a symmetry detection and analysis scheme for sets of points in IRn and its extension to image analysis by way of local features. Each object is represented by a set of points S 2 IRn, where the symmetry is manifested by the multiple self-alignments of S. The alignment problem is formulated as a quadratic binary optimization problem, with an efficient solution via spectral relaxation. For symmetric objects, this results in a multiplicity of eigenvalues whose corresponding eigenvectors allow the detection and analysis of both types of symmetry. We improve the scheme's robustness by incorporating geometrical constraints into the spectral analysis. Our approach is experimentally verified by applying it to 2D and 3D synthetic objects as well as real images.
Michael Chertok, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Efficient High Order Matching
abstract
We present a computational approach to high-order matching of data sets in IR(d). Those are matchings based on data affinity measures that score the matching of more than two pairs of points at a time. High-order affinities are represented by tensors and the matching is then given by a rank-one approximation of the affinity tensor and a corresponding discretization. Our approach is rigorously justified by extending Zass and Shashua's hypergraph matching to high-order spectral matching. This paves the way for a computationally efficient dual-marginalization spectral matching scheme. We also show that, based on the spectral properties of random matrices, affinity tensors can be randomly sparsified while retaining the matching accuracy. Our contributions are experimentally validated by applying them to synthetic as well as real data sets.
Michael Chertok, Yosi Keller
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Improving Shape Retrieval by Spectral Matching and Meta Similarity
abstract
We propose two computational approaches for improving the retrieval of planar shapes. First, we suggest a geometrically motivated quadratic similarity measure, that is optimized by way of spectral relaxation of a quadratic assignment. By utilizing state-of-the-art shape descriptors and a pairwise serialization constraint, we derive a formulation that is resilient to boundary noise, articulations and nonrigid deformations. This allows both shape matching and retrieval. We also introduce a shape meta-similarity measure that agglomerates pairwise shape similarities and improves the retrieval accuracy. When applied to the MPEG-7 shape dataset in conjunction with the proposed geometric matching scheme, we obtained a retrieval rate of 92.5%.
Amir Egozi, Yosi Keller, Hugo Guterman
IEEE Trans. Image Process.2
2008 Global parametric image alignment via high-order approximation
Yosi Keller, Amir Averbuch
Comput. Vis. Image Underst.1
2007 A projection-based extension to phase correlation image alignment
Yosi Keller, Amir Averbuch
Signal Process.1
2006 Multisensor Image Registration via Implicit Similarity
abstract
This paper presents an approach to the registration of significantly dissimilar images, acquired by sensors of different modalities. A robust matching criterion is derived by aligning the locations of gradient maxima. The alignment is achieved by iteratively maximizing the magnitudes of the intensity gradients of a set of pixels in one of the images, where the set is initialized by the gradient maxima locations of the second image. No explicit similarity measure that uses the intensities of both images is used. The computation utilizes the full spatial information of the first image and the accuracy and robustness of the registration depend only on it. False matchings are detected and adaptively weighted using a directional similarity measure. By embedding the scheme in a "coarse to fine" formulation, we were able to estimate affine and projective global motions, even when the images were characterized by complex space varying intensity transformations. The scheme is especially suitable when one of the images is of considerably better quality than the other (noise, blur, etc.). We demonstrate these properties via experiments on real multisensor image sets.
Yosi Keller, Amir Averbuch
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Data Fusion and Multicue Data Matching by Diffusion Maps
abstract
Data fusion and multicue data matching are fundamental tasks of high-dimensional data analysis. In this paper, we apply the recently introduced diffusion framework to address these tasks. Our contribution is three-fold: First, we present the Laplace-Beltrami approach for computing density invariant embeddings which are essential for integrating different sources of data. Second, we describe a refinement of the Nyström extension algorithm called "geometric harmonics." We also explain how to use this tool for data assimilation. Finally, we introduce a multicue data matching scheme based on nonlinear spectral graphs alignment. The effectiveness of the presented schemes is validated by applying it to the problems of lipreading and image sequence alignment.
Stéphane Lafon, Yosi Keller, Ronald R. Coifman
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 A signal processing approach to symmetry detection
abstract
We present an algorithm that detects rotational and reflectional symmetries of two-dimensional objects. Both symmetry types are effectively detected and analyzed using the angular correlation (AC), which measures the correlation between images in the angular direction. The AC is accurately computed using the pseudopolar Fourier transform, which rapidly computes the Fourier transform of an image on a near-polar grid. We prove that the AC of symmetric images is a periodic signal whose frequency is related to the order of the symmetry. This frequency is recovered via spectrum estimation, which is a proven technique in signal processing with a variety of efficient solutions. We also provide a novel approach for finding the center of symmetry and demonstrate the applicability of our scheme to the analysis of real images.
Yosi Keller, Yoel Shkolnisky
IEEE Trans. Image Process.1
2005 Robust Image Alignment using Third-order Global Motion Estimation
Yosi Keller, Amir Averbuch
BMVC1
2005 Algebraically Accurate Volume Registration Using Euler's Theorem and the 3-D Pseudo-Polar FFT
abstract
We present an algorithm for the registration of rotated and translated volumes, which operates in the frequency domain. The Fourier domain allows to compute the rotation and translation parameters separately, thus reducing a problem with six degrees of freedom to two problems of three degrees of freedom each. We propose a three-step procedure. The first step estimates the rotation axis. The second computes the planar rotation relative to the rotation axis, and the third recovers the translational displacement by using the phase correlation technique. The rotation estimation is based on Euler's theorem, which allows to represent a rotation using only three parameters. Two parameters represent the rotation axis and one parameter represents the planar rotation perpendicular to the axis. By using the 3D pseudo-polar FFT, the estimation of the rotation axis is shown to be algebraically accurate. A variant of the angular difference function registration algorithm is derived for the estimation of the planar rotation around the axis. The experimental results show that the algorithm is accurate and robust to noise.
Yosi Keller, Amir Averbuch, Yoel Shkolnisky
CVPR (2)1
2005 A non Cartesian FFT approach to image alignment
abstract
The estimation of large motions without prior knowledge is an important problem in image registration. In this paper we present the angular difference function (ADF) and demonstrate its applicability to rotation estimation. The ADF of two functions is defined as the integral of their spectral difference along the radial direction. It is efficiently computed using the pseudo-polar Fourier transform, which computes the discrete Fourier transform of an image on a near spherical grid. Unlike other Fourier based registration schemes, the suggested approach does not require any interpolation. Thus, it is more accurate and significantly faster.
Yosi Keller, Yoel Shkolnisky, Amir Averbuch
ICIP (3)1
2005 Accurate multi-dimensional alignment
abstract
We present an algorithm for aligning rotated and translated volumes, which operates in the frequency domain. The Fourier domain allows us to compute the rotation and translation parameters separately, thus reducing a problem with six degrees of freedom to two problems of three degrees of freedom each. We propose a three-step procedure. The first step estimates the rotation axis, the second computes the planar rotation relative to the rotation axis, and the third recovers the translational displacement by using the phase correlation technique. By using the 3-D pseudo-polar FFT, the estimation of the rotation axis is shown to be algebraically accurate. Experimental results show that the algorithm is accurate and robust to noise.
Yosi Keller, Yoel Shkolnisky, Amir Averbuch
ICIP (3)1
2005 The Angular Difference Function and Its Application to Image Registration
abstract
The estimation of large motions without prior knowledge is an important problem in image registration. In this paper, we present the angular difference function (ADF) and demonstrate its applicability to rotation estimation. The ADF of two functions is defined as the integral of their spectral difference along the radial direction. It is efficiently computed using the pseudopolar Fourier transform, which computes the discrete Fourier transform of an image on a near spherical grid. Unlike other Fourier-based registration schemes, the suggested approach does not require any interpolation. Thus, it is more accurate and significantly faster.
Yosi Keller, Yoel Shkolnisky, Amir Averbuch
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Pseudopolar-based estimation of large translations, rotations, and scalings in images
abstract
One of the major challenges related to image registration is the estimation of large motions without prior knowledge. This paper presents a Fourier-based approach that estimates large translations, scalings, and rotations. The algorithm uses the pseudopolar (PP) Fourier transform to achieve substantial improved approximations of the polar and log-polar Fourier transforms of an image. Thus, rotations and scalings are reduced to translations which are estimated using phase correlation. By utilizing the PP grid, we increase the performance (accuracy, speed, and robustness) of the registration algorithms. Scales up to 4 and arbitrary rotation angles can be robustly recovered, compared to a maximum scaling of 2 recovered by state-of-the-art algorithms. The algorithm only utilizes one-dimensional fast Fourier transform computations whose overall complexity is significantly lower than prior works. Experimental results demonstrate the applicability of the proposed algorithms.
Yosi Keller, Amir Averbuch, Moshe Israeli
IEEE Trans. Image Process.1
2004 Fast motion estimation using bidirectional gradient methods
abstract
Gradient-based motion estimation methods (GMs) are considered to be in the heart of state-of-the-art registration algorithms, being able to account for both pixel and subpixel registration and to handle various motion models (translation, rotation, affine, and projective). These methods estimate the motion between two images based on the local changes in the image intensities while assuming image smoothness. This paper offers two main contributions. The first is enhancement of the GM technique by introducing two new bidirectional formulations of the GM. These improve the convergence properties for large motions. The second is that we present an analytical convergence analysis of the GM and its properties. Experimental results demonstrate the applicability of these algorithms to real images.
Yosi Keller, Amir Averbuch
IEEE Trans. Image Process.1
2003 Implicit similarity: a new approach to multi-sensor image registration
abstract
This paper presents an implicit similarity-based approach to registration of significantly dissimilar images, acquired by sensors at different modalities. The proposed algorithm introduces a robust matching criterion by aligning the locations of gradient maxima. The alignment is formulated as a parametric variational optimization problem, which is solved iteratively by considering the intensities of a single image. The location of the maxima of the second image's gradient are used as initialization. We are able to robustly estimate affine and projective global motions using 'coarse to fine' processing, even when the images are characterized by complex space varying intensity transformations. Finally, we present the registration of real images, which were taken by multi-sensor and multi-modality using affine and projective motion models.
Yosi Keller, Amir Averbuch
CVPR (2)1
2003 Fast gradient methods based on global motion estimation for video compression
abstract
This paper presents a fast global motion estimation (GME) algorithm based on gradient methods (GM), which can be used for real-time applications, such as in MPEG4 video compression. This approach improves the existing state-of-the-art GME algorithms by introducing two major modifications: first, only a small subset (down to 3%) of the original image pixels is used in the estimation process. Second, an interpolation-free formulation of the basic GM is derived, further decreasing the computational complexity. Experimental results show no loss of GME accuracy and compression efficiency compared to the MPEG-4 verification model, while reducing the computational complexity of the GME by a factor of 20.
Yosi Keller, Amir Averbuch
IEEE Trans. Circuits Syst. Video Technol.1
2002 FFT based image registration
abstract
We present a new unified approach to FFT based image registration. Prior works divided the registration process into two stages: the first was based on phase correlation (PC) which provides pixel accurate registration [5], while the second step provides subpixel registration accuracy [1, 3]. By extending the PC method we derive a FFT based image registration algorithm which is able to estimate large translations with subpixel accuracy. The algorithm's properties resemble those of the Gradient Methods [4] while outperforming it by exhibiting superior convergence range.
Amir Averbuch, Yosi Keller
ICASSP2
2002 Fast motion estimation using Bidirectional Gradient Methods
abstract
Gradient based motion estimation techniques (GM) are considered to be in the heart of state-of-the-art registration algorithms [2], being able to account for both pixel and subpixel registration and to handle various motion models (translation, rotation, affine, projective). These methods estimate the motion between two images based on the local changes in the image intensities while assuming image smoothness. This paper introduces two new bidirectional formulations of the GM, improving the convergence properties for large motions. Experimental results demonstrate the applicability of these algorithms to real images.
Amir Averbuch, Yosi Keller
ICASSP2
2002 Fast gradient methods based global motion estimation for video compression
abstract
This paper presents a fast global motion estimation (GME) algorithm based on gradient methods (GM), which can be used for real-time applications, such as MPEG-4 video compression. Our approach improves existing state-of-the-art GME algorithms by introducing two major modifications: first, only a small subset (down-to 3%) of the original image pixels is used in the estimation process. Second, a warp-free formulation of the basic GM is derived, further decreasing the computational complexity. Experimental results show that a 20 fold computation complexity reduction is achieved, without compromising the GME accuracy and compression efficiency.
Yosi Keller, Amir Averbuch, Ofer Miller
ICIP (1)1