VLDB 2026 Research / reviewers in the wild / expert
Ronen Basri
dblp:b/RonenBasri
· DBLP profile ↗
112ranked-venue papers
35as first author
9since 2021 · last 2025
0000-0001-8053-2151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 35 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 16 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RESfM: Robust Deep Equivariant Structure from MotionabstractMultiview Structure from Motion is a fundamental and challenging computer vision problem. A recent deep-based approach utilized matrix equivariant architectures for simultaneous recovery of camera pose and 3D scene structure from large image collections. That work, however, made the unrealistic assumption that the point tracks given as input are almost clean of outliers. Here, we propose an architecture suited to dealing with outliers by adding a multiview inlier/outlier classification module that respects the model equivariance and by utilizing a robust bundle adjustment step. Experiments demonstrate that our method can be applied successfully in realistic settings that include large image collections and point tracks extracted with common heuristics that include many outliers, achieving state-of-the-art accuracies in almost all runs, superior to existing deep-based methods and on-par with leading classical (non-deep) sequential and global methods. Fadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun, Ronen Basri |
ICLR | 5 |
| 2024 | Likelihood Training of Cascaded Diffusion Models via Hierarchical Volume-preserving MapsabstractCascaded models are multi-scale generative models with a marked capacity for producing perceptually impressive samples at high resolutions. In this work, we show that they can also be excellent likelihood models, so long as we overcome a fundamental difficulty with probabilistic multi-scale models: the intractability of the likelihood function. Chiefly, in cascaded models each intermediary scale introduces extraneous variables that cannot be tractably marginalized out for likelihood evaluation. This issue vanishes by modeling the diffusion process on latent spaces induced by a class of transformations we call hierarchical volume-preserving maps, which decompose spatially structured data in a hierarchical fashion without introducing local distortions in the latent space. We demonstrate that two such maps are well-known in the literature for multiscale modeling: Laplacian pyramids and wavelet transforms. Not only do such reparameterizations allow the likelihood function to be directly expressed as a joint likelihood over the scales, we show that the Laplacian pyramid and wavelet transform also produces significant improvements to the state-of-the-art on a selection of benchmarks in likelihood modeling, including density estimation, lossless compression, and out-of-distribution detection. Investigating the theoretical basis of our empirical gains we uncover deep connections to score matching under the Earth Mover's Distance (EMD), which is a well-known surrogate for perceptual similarity. Henry Li, Ronen Basri, Yuval Kluger |
ICLR | 2 |
| 2024 | Consensus Learning with Deep Sets for Essential Matrix EstimationabstractRobust estimation of the essential matrix, which encodes the relative position and orientation of two cameras, is a fundamental step in structure from motion pipelines. Recent deep-based methods achieved accurate estimation by using complex network architectures that involve graphs, attention layers, and hard pruning steps. Here, we propose a simpler network architecture based on Deep Sets. Given a collection of point matches extracted from two images, our method identifies outlier point matches and models the displacement noise in inlier matches. A weighted DLT module uses these predictions to regress the essential matrix. Our network achieves accurate recovery that is superior to existing networks with significantly more complex architectures. Dror Moran, Yuval Margalit, Guy Trostianetsky, Fadi Khatib, Meirav Galun, Ronen Basri |
NeurIPS | 6 |
| 2024 | CALVIN: Improved Contextual Video Captioning via Instruction TuningabstractThe recent emergence of powerful Vision-Language models (VLMs) has significantly improved image captioning. Some of these models are extended to caption videos as well. However, their capabilities to understand complex scenes are limited, and the descriptions they provide for scenes tend to be overly verbose and focused on the superficial appearance of objects. Scene descriptions, especially in movies, require a deeper contextual understanding, unlike general-purpose video captioning. To address this challenge, we propose a model, CALVIN, a specialized video LLM that leverages previous movie context to generate fully "contextual" scene descriptions. To achieve this, we train our model on a suite of tasks that integrate both image-based question-answering and video captioning within a unified framework, before applying instruction tuning to refine the model's ability to provide scene captions. Lastly, we observe that our model responds well to prompt engineering and few-shot in-context learning techniques, enabling the user to adapt it to any new movie with very little additional annotation. Gowthami Somepalli, Arkabandhu Chowdhury, Jonas Geiping, Ronen Basri, Tom Goldstein, David Jacobs 0001 |
NeurIPS | 4 |
| 2024 | Spectral Analysis of the Neural Tangent Kernel for Deep Residual NetworksabstractDeep residual network architectures have been shown to achieve superior accuracy over classical feed-forward networks, yet their success is still not fully understood. Focusing on massively over-parameterized, fully connected residual networks with ReLU activation through their respective neural tangent kernels (ResNTK), we provide here a spectral analysis of these kernels. Specifically, we show that, much like NTK for fully connected networks (FC-NTK), for input distributed uniformly on the hypersphere $S^d$, the eigenvalues of ResNTK corresponding to their spherical harmonics eigenfunctions decay polynomially with frequency $k$ as $k^{-d}$. These in turn imply that the set of functions in their Reproducing Kernel Hilbert Space are identical to those of both FC-NTK as well as the standard Laplace kernel. Our spectral analysis allows us to highlight several additional properties of ResNTK, which depend on the choice of a hyper-parameter that balances between the skip and residual connections. Specifically, (1) with no bias, deep ResNTK is significantly biased toward even frequency functions; (2) unlike FC-NTK for deep networks, which is spiky and therefore yields poor generalization, ResNTK is stable and yields small generalization errors. We finally demonstrate these with experiments showing further that these phenomena arise in real networks. Yuval Belfer, Amnon Geifman, Meirav Galun, Ronen Basri |
J. Mach. Learn. Res. | 4 |
| 2023 | A Kernel Perspective of Skip Connections in Convolutional Networks
Daniel Barzilai, Amnon Geifman, Meirav Galun, Ronen Basri |
ICLR | 4 |
| 2022 | On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process KernelsabstractWe study the properties of various over-parameterized convolutional neural architectures through their respective Gaussian Process and Neural Tangent kernels. We prove that, with normalized multi-channel input and ReLU activation, the eigenfunctions of these kernels with the uniform measure are formed by products of spherical harmonics, defined over the channels of the different pixels. We next use hierarchical factorizable kernels to bound their respective eigenvalues. We show that the eigenvalues decay polynomially, quantify the rate of decay, and derive measures that reflect the composition of hierarchical features in these networks. Our theory provides a concrete quantitative characterization of the role of locality and hierarchy in the inductive bias of over-parameterized convolutional network architectures. Amnon Geifman, Meirav Galun, David Jacobs 0001, Ronen Basri |
NeurIPS | 4 |
| 2021 | Deep Permutation Equivariant Structure from MotionabstractExisting deep methods produce highly accurate 3D reconstructions in stereo and multiview stereo settings, i.e., when cameras are both internally and externally calibrated. Nevertheless, the challenge of simultaneous recovery of camera poses and 3D scene structure in multiview settings with deep networks is still outstanding. Inspired by projective factorization for Structure from Motion (SFM) and by deep matrix completion techniques, we propose a neural network architecture that, given a set of point tracks in multiple images of a static scene, recovers both the camera parameters and a (sparse) scene structure by minimizing an unsupervised reprojection loss. Our network architecture is designed to respect the structure of the problem: the sought output is equivariant to permutations of both cameras and scene points. Notably, our method does not require initialization of camera parameters or 3D point locations. We test our architecture in two setups: (1) single scene reconstruction and (2) learning from multiple scenes. Our experiments, conducted on a variety of datasets in both internally calibrated and uncalibrated settings, indicate that our method accurately recovers pose and structure, on par with classical state of the art methods. Additionally, we show that a pre-trained network can be used to reconstruct novel scenes using inexpensive fine-tuning with no loss of accuracy. Dror Moran, Hodaya Koslowsky, Yoni Kasten, Haggai Maron, Meirav Galun, Ronen Basri |
ICCV | 6 |
| 2021 | Shift Invariance Can Reduce Adversarial RobustnessabstractShift invariance is a critical property of CNNs that improves performance on classification. However, we show that invariance to circular shifts can also lead to greater sensitivity to adversarial attacks. We first characterize the margin between classes when a shift-invariant {\em linear} classifier is used. We show that the margin can only depend on the DC component of the signals. Then, using results about infinitely wide networks, we show that in some simple cases, fully connected and shift-invariant neural networks produce linear decision boundaries. Using this, we prove that shift invariance in neural networks produces adversarial examples for the simple case of two classes, each consisting of a single image with a black or white dot on a gray background. This is more than a curiosity; we show empirically that with real datasets and realistic architectures, shift invariance reduces adversarial robustness. Finally, we describe initial experiments using synthetic data to probe the source of this connection. Vasu Singla, Songwei Ge, Ronen Basri, David Jacobs 0001 |
NeurIPS | 3 |
| 2020 | Averaging Essential and Fundamental Matrices in Collinear Camera SettingsabstractGlobal methods to Structure from Motion have gained popularity in recent years. A significant drawback of global methods is their sensitivity to collinear camera settings. In this paper, we introduce an analysis and algorithms for averaging bifocal tensors (essential or fundamental matrices) when either subsets or all of the camera centers are collinear. We provide a complete spectral characterization of bifocal tensors in collinear scenarios and further propose two averaging algorithms. The first algorithm uses rank constrained minimization to recover camera matrices in fully collinear settings. The second algorithm enriches the set of possibly mixed collinear and non-collinear cameras with additional, ``virtual cameras," which are placed in general position, enabling the application of existing averaging methods to the enriched set of bifocal tensors. Our algorithms are shown to achieve state of the art results on various benchmarks that include autonomous car datasets and unordered image collections in both calibrated and unclibrated settings. Amnon Geifman, Yoni Kasten, Meirav Galun, Ronen Basri |
CVPR | 4 |
| 2020 | Frequency Bias in Neural Networks for Input of Non-Uniform DensityabstractRecent works have partly attributed the generalization ability of over-parameterized neural networks to frequency bias – networks trained with gradient descent on data drawn from a uniform distribution find a low frequency fit before high frequency ones. As realistic training sets are not drawn from a uniform distribution, we here use the Neural Tangent Kernel (NTK) model to explore the effect of variable density on training dynamics. Our results, which combine analytic and empirical observations, show that when learning a pure harmonic function of frequency $\kappa$, convergence at a point $x \in \S^{d-1}$ occurs in time $O(\kappa^d/p(x))$ where $p(x)$ denotes the local density at $x$. Specifically, for data in $\S^1$ we analytically derive the eigenfunctions of the kernel associated with the NTK for two-layer networks. We further prove convergence results for deep, fully connected networks with respect to the spectral decomposition of the NTK. Our empirical study highlights similarities and differences between deep and shallow networks in this model. Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs 0001, Yoni Kasten, Shira Kritchman |
ICML | 1 |
| 2020 | Learning Algebraic Multigrid Using Graph Neural NetworksabstractEfficient numerical solvers for sparse linear systems are crucial in science and engineering. One of the fastest methods for solving large-scale sparse linear systems is algebraic multigrid (AMG). The main challenge in the construction of AMG algorithms is the selection of the prolongation operator—a problem-dependent sparse matrix which governs the multiscale hierarchy of the solver and is critical to its efficiency. Over many years, numerous methods have been developed for this task, and yet there is no known single right answer except in very special cases. Here we propose a framework for learning AMG prolongation operators for linear systems with sparse symmetric positive (semi-) definite matrices. We train a single graph neural network to learn a mapping from an entire class of such matrices to prolongation operators, using an efficient unsupervised loss function. Experiments on a broad class of problems demonstrate improved convergence rates compared to classical AMG, demonstrating the potential utility of neural networks for developing sparse system solvers. Ilay Luz, Meirav Galun, Haggai Maron, Ronen Basri, Irad Yavneh |
ICML | 4 |
| 2020 | On the Similarity between the Laplace and Neural Tangent KernelsabstractRecent theoretical work has shown that massively overparameterized neural networks are equivalent to kernel regressors that use Neural Tangent Kernels (NTKs). Experiments show that these kernel methods perform similarly to real neural networks. Here we show that NTK for fully connected networks with ReLU activation is closely related to the standard Laplace kernel. We show theoretically that for normalized data on the hypersphere both kernels have the same eigenfunctions and their eigenvalues decay polynomially at the same rate, implying that their Reproducing Kernel Hilbert Spaces (RKHS) include the same sets of functions. This means that both kernels give rise to classes of functions with the same smoothness properties. The two kernels differ for data off the hypersphere, but experiments indicate that when data is properly normalized these differences are not significant. Finally, we provide experiments on real data comparing NTK and the Laplace kernel, along with a larger class of $\gamma$-exponential kernels. We show that these perform almost identically. Our results suggest that much insight about neural networks can be obtained from analysis of the well-known Laplace kernel, which has a simple closed form. Amnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun, David Jacobs 0001, Ronen Basri |
NeurIPS | 6 |
| 2020 | Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceabstractIn this work we address the challenging problem of multiview 3D surface reconstruction. We introduce a neural network architecture that simultaneously learns the unknown geometry, camera parameters, and a neural renderer that approximates the light reflected from the surface towards the camera. The geometry is represented as a zero level-set of a neural network, while the neural renderer, derived from the rendering equation, is capable of (implicitly) modeling a wide set of lighting conditions and materials. We trained our network on real world 2D images of objects with different material properties, lighting conditions, and noisy camera initializations from the DTU MVS dataset. We found our model to produce state of the art 3D surface reconstructions with high fidelity, resolution and detail. Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Ronen Basri, Yaron Lipman |
NeurIPS | 6 |
| 2020 | On Detection of Faint Edges in Noisy ImagesabstractA fundamental question for edge detection in noisy images is how faint can an edge be and still be detected. In this paper we offer a formalism to study this question and subsequently introduce computationally efficient multiscale edge detection algorithms designed to detect faint edges in noisy images. In our formalism we view edge detection as a search in a discrete, though potentially large, set of feasible curves. First, we derive approximate expressions for the detection threshold as a function of curve length and the complexity of the search space. We then present two edge detection algorithms, one for straight edges, and the second for curved ones. Both algorithms efficiently search for edges in a large set of candidates by hierarchically constructing difference filters that match the curves traced by the sought edges. We demonstrate the utility of our algorithms in both simulations and applications involving challenging real images. Finally, based on these principles, we develop an algorithm for fiber detection and enhancement. We exemplify its utility to reveal and enhance nerve axons in light microscopy images. Nati Ofir, Meirav Galun, Sharon Alpert, Achi Brandt, Boaz Nadler, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | GPSfM: Global Projective SFM Using Algebraic Constraints on Multi-View Fundamental MatricesabstractThis paper addresses the problem of recovering projective camera matrices from collections of fundamental matrices in multiview settings. We make two main contributions. First, given (2n) fundamental matrices computed for n images, we provide a complete algebraic characterization in the form of conditions that are both necessary and sufficient to enabling the recovery of camera matrices. These conditions are based on arranging the fundamental matrices as blocks in a single matrix, called the n-view fundamental matrix, and characterizing this matrix in terms of the signs of its eigenvalues and rank structures. Secondly, we propose a concrete algorithm for projective structure-formation that utilizes this characterization. Given a complete or partial collection of measured-fundamental matrices, our method seeks camera matrices that minimize a global algebraic error for the measured fundamental matrices. In contrast to existing methods, our optimization, without any initialization, produces a consistent set of fundamental matrices that corresponds to a unique set of cameras (up to a choice of projective frame). Our experiments indicate that our method achieves state of the art performance in both accuracy and running time. Yoni Kasten, Amnon Geifman, Meirav Galun, Ronen Basri |
CVPR | 4 |
| 2019 | Algebraic Characterization of Essential Matrices and Their Averaging in Multiview SettingsabstractEssential matrix averaging, i.e., the task of recovering camera locations and orientations in calibrated, multiview settings, is a first step in global approaches to Euclidean structure from motion. A common approach to essential matrix averaging is to separately solve for camera orientations and subsequently for camera positions. This paper presents a novel approach that solves simultaneously for both camera orientations and positions. We offer a complete characterization of the algebraic conditions that enable a unique Euclidean reconstruction of n cameras from a collection of (2n) essential matrices. We next use these conditions to formulate essential matrix averaging as a constrained optimization problem, allowing us to recover a consistent set of essential matrices given a (possibly partial) set of measured essential matrices computed independently for pairs of images. We finally use the recovered essential matrices to determine the global positions and orientations of the n cameras. We test our method on common SfM datasets, demonstrating high accuracy while maintaining efficiency and robustness, compared to existing methods. Yoni Kasten, Amnon Geifman, Meirav Galun, Ronen Basri |
ICCV | 4 |
| 2019 | Learning to Optimize Multigrid PDE SolversabstractConstructing fast numerical solvers for partial differential equations (PDEs) is crucial for many scientific disciplines. A leading technique for solving large-scale PDEs is using multigrid methods. At the core of a multigrid solver is the prolongation matrix, which relates between different scales of the problem. This matrix is strongly problem-dependent, and its optimal construction is critical to the efficiency of the solver. In practice, however, devising multigrid algorithms for new problems often poses formidable challenges. In this paper we propose a framework for learning multigrid solvers. Our method learns a (single) mapping from discretized PDEs to prolongation operators for a broad class of 2D diffusion problems. We train a neural network once for the entire class of PDEs, using an efficient and unsupervised loss function. Our tests demonstrate improved convergence rates compared to the widely used Black-Box multigrid scheme, suggesting that our method successfully learned rules for constructing prolongation matrices. Daniel Greenfeld, Meirav Galun, Ronen Basri, Irad Yavneh, Ron Kimmel |
ICML | 3 |
| 2019 | The Convergence Rate of Neural Networks for Learned Functions of Different FrequenciesabstractWe study the relationship between the frequency of a function and the speed at which a neural network learns it. We build on recent results that show that the dynamics of overparameterized neural networks trained with gradient descent can be well approximated by a linear system. When normalized training data is uniformly distributed on a hypersphere, the eigenfunctions of this linear system are spherical harmonic functions. We derive the corresponding eigenvalues for each frequency after introducing a bias term in the model. This bias term had been omitted from the linear network model without significantly affecting previous theoretical results. However, we show theoretically and experimentally that a shallow neural network without bias cannot represent or learn simple, low frequency functions with odd frequencies. Our results lead to specific predictions of the time it will take a network to learn functions of varying frequency. These predictions match the empirical behavior of both shallow and deep networks. Ronen Basri, David Jacobs 0001, Yoni Kasten, Shira Kritchman |
NeurIPS | 1 |
| 2019 | Resultant Based Incremental Recovery of Camera Pose From Pairwise MatchesabstractIncremental (online) structure from motion pipelines seek to recover the camera matrix associated with an image I_n given n-1 images, I_1,...,I_n-1, whose camera matrices have already been recovered. In this paper, we introduce a novel solution to the six-point online algorithm to recover the exterior parameters associated with I_n. Our algorithm uses just six corresponding pairs of 2D points, extracted each from I_n and from any of the preceding n-1 images, allowing the recovery of the full six degrees of freedom of the n'th camera, and unlike common methods, does not require tracking feature points in three or more images. Our novel solution is based on constructing a Dixon resultant, yielding a solution method that is both efficient and accurate compared to existing solutions. We further use Bernstein's theorem to prove a tight bound on the number of complex solutions. Our experiments demonstrate the utility of our approach. Yoni Kasten, Meirav Galun, Ronen Basri |
WACV | 3 |
| 2018 | SpectralNet: Spectral Clustering using Deep Neural Networks
Uri Shaham 0001, Kelly P. Stanton, Henry Li, Ronen Basri, Boaz Nadler, Yuval Kluger |
ICLR (Poster) | 4 |
| 2018 | Elasticity-based matching by minimising the symmetric difference of shapesabstractThe authors consider the problem of matching two shapes assuming these shapes are related by an elastic deformation. Using linearised elasticity theory and the finite‐element method, they seek an elastic deformation that is caused by simple external boundary forces and accounts for the difference between the two shapes. The main contribution is in proposing a cost function and an optimisation procedure to minimise the symmetric difference between the deformed and the target shapes as an alternative to point matches that guide the matching in other techniques. The authors show how to approximate the non‐linear optimisation problem by a sequence of convex problems. They demonstrate the utility of the proposed method in experiments and compare it to an iterative closest point like matching algorithm. Konrad Simon, Ronen Basri |
IET Comput. Vis. | 2 |
| 2017 | A New Rank Constraint on Multi-view Fundamental Matrices, and Its Application to Camera Location RecoveryabstractAccurate estimation of camera matrices is an important step in structure from motion algorithms. In this paper we introduce a novel rank constraint on collections of fundamental matrices in multi-view settings. We show that in general, with the selection of proper scale factors, a matrix formed by stacking fundamental matrices between pairs of images has rank 6. Moreover, this matrix forms the symmetric part of a rank 3 matrix whose factors relate directly to the corresponding camera matrices. We use this new characterization to produce better estimations of fundamental matrices by optimizing an L1-cost function using Iterative Re-weighted Least Squares and Alternate Direction Method of Multiplier. We further show that this procedure can improve the recovery of camera locations, particularly in multi-view settings in which fewer images are available. Roni Sengupta, Tal Amir, Meirav Galun, Tom Goldstein, David Jacobs 0001, Amit Singer, Ronen Basri |
CVPR | 7 |
| 2017 | Efficient Representation of Low-Dimensional Manifolds using Deep Networks
Ronen Basri, David Jacobs 0001 |
ICLR (Poster) | 1 |
| 2016 | Fast Detection of Curved Edges at Low SNRabstractDetecting edges is a fundamental problem in computer vision with many applications, some involving very noisy images. While most edge detection methods are fast, they perform well only on relatively clean images. Unfortunately, sophisticated methods that are robust to high levels of noise are quite slow. In this paper we develop a novel multiscale method to detect curved edges in noisy images. Even though our algorithm searches for edges over an exponentially large set of candidate curves, its runtime is nearly linear in the total number of image pixels. As we demonstrate experimentally, our algorithm is orders of magnitude faster than previous methods designed to deal with high noise levels. At the same time it obtains comparable and often superior results to existing methods on a variety of challenging noisy images. Nati Ofir, Meirav Galun, Boaz Nadler, Ronen Basri |
CVPR | 4 |
| 2016 | Learning 3D Deformation of Animals from 2D ImagesabstractAbstract Understanding how an animal can deform and articulate is essential for a realistic modification of its 3D model. In this paper, we show that such information can be learned from user‐clicked 2D images and a template 3D model of the target animal. We present a volumetric deformation framework that produces a set of new 3D models by deforming a template 3D model according to a set of user‐clicked images. Our framework is based on a novel locally‐bounded deformation energy, where every local region has its own stiffness value that bounds how much distortion is allowed at that location. We jointly learn the local stiffness bounds as we deform the template 3D mesh to match each user‐clicked image. We show that this seemingly complex task can be solved as a sequence of convex optimization problems. We demonstrate the effectiveness of our approach on cats and horses, which are highly deformable and articulated animals. Our framework produces new 3D models of animals that are significantly more plausible than methods without learned stiffness. Angjoo Kanazawa, Shahar Z. Kovalsky, Ronen Basri, David Jacobs 0001 |
Comput. Graph. Forum | 3 |
| 2016 | Guest Editorial: Special Section on CVPR 2014abstractThe papers in this special section were presented at the IEEE Computer Vision and Pattern Recognition (CVPR), June, 2014, jointly sponsored by the IEEE and the Computer Vision Foundation. Ronen Basri, Cornelia Fermüller, Aleix Martinez, René Vidal |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Wide Baseline Stereo Matching with Convex Bounded Distortion ConstraintsabstractFinding correspondences in wide baseline setups is a challenging problem. Existing approaches have focused largely on developing better feature descriptors for correspondence and on accurate recovery of epipolar line constraints. This paper focuses on the challenging problem of finding correspondences once approximate epipolar constraints are given. We introduce a novel method that integrates a deformation model. Specifically, we formulate the problem as finding the largest number of corresponding points related by a bounded distortion map that obeys the given epipolar constraints. We show that, while the set of bounded distortion maps is not convex, the subset of maps that obey the epipolar line constraints is convex, allowing us to introduce an efficient algorithm for matching. We further utilize a robust cost function for matching and employ majorization-minimization for its optimization. Our experiments indicate that our method finds significantly more accurate maps than existing approaches. Meirav Galun, Tal Amir, Tal Hassner, Ronen Basri, Yaron Lipman |
ICCV | 4 |
| 2015 | A Multiscale Variable-Grouping Framework for MRF Energy MinimizationabstractWe present a multiscale approach for minimizing the energy associated with Markov Random Fields (MRFs) with energy functions that include arbitrary pairwise potentials. The MRF is represented on a hierarchy of successively coarser scales, where the problem on each scale is itself an MRF with suitably defined potentials. These representations are used to construct an efficient multiscale algorithm that seeks a minimal-energy solution to the original problem. The algorithm is iterative and features a bidirectional crosstalk between fine and coarse representations. We use consistency criteria to guarantee that the energy is nonincreasing throughout the iterative process. The algorithm is evaluated on real-world datasets, achieving competitive performance in relatively short run-times. Omer Meir, Meirav Galun, Stav Yagev, Ronen Basri, Irad Yavneh |
ICCV | 4 |
| 2015 | Tight Relaxation of Quadratic MatchingabstractAbstract Establishing point correspondences between shapes is extremely challenging as it involves both finding sets of semantically persistent feature points, as well as their combinatorial matching. We focus on the latter and consider the Quadratic Assignment Matching (QAM) model. We suggest a novel convex relaxation for this NP‐hard problem that builds upon a rank‐one reformulation of the problem in a higher dimension, followed by relaxation into a semidefinite program (SDP). Our method is shown to be a certain hybrid of the popular spectral and doubly‐stochastic relaxations of QAM and in particular we prove that it is tighter than both. Experimental evaluation shows that the proposed relaxation is extremely tight: in the majority of our experiments it achieved the certified global optimum solution for the problem, while other relaxations tend to produce sub‐optimal solutions. This, however, comes at the price of solving an SDP in a higher dimension. Our approach is further generalized to the problem of Consistent Collection Matching (CCM), where we solve the QAM on a collection of shapes while simultaneously incorporating a global consistency constraint. Lastly, we demonstrate an application to metric learning of collections of shapes. Itay Kezurer, Shahar Z. Kovalsky, Ronen Basri, Yaron Lipman |
Comput. Graph. Forum | 3 |
| 2015 | From Shading to Local ShapeabstractWe develop a framework for extracting a concise representation of the shape information available from diffuse shading in a small image patch. This produces a mid-level scene descriptor, comprised of local shape distributions that are inferred separately at every image patch across multiple scales. The framework is based on a quadratic representation of local shape that, in the absence of noise, has guarantees on recovering accurate local shape and lighting. And when noise is present, the inferred local shape distributions provide useful shape information without over-committing to any particular image explanation. These local shape distributions naturally encode the fact that some smooth diffuse regions are more informative than others, and they enable efficient and robust reconstruction of object-scale shape. Experimental results show that this approach to surface reconstruction compares well against the state-of-art on both synthetic images and captured photographs. Ayan Chakrabarti, Ronen Basri, Steven J. Gortler, David Jacobs 0001, Todd E. Zickler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Detection of Long Edges on a Computational Budget: A Sublinear ApproachabstractEdge detection is a challenging, important task in image analysis. Various applications require real-time detection of long edges in large and noisy images, possibly under limited computational resources. While standard edge detection methods are computationally fast, they perform well only at low levels of noise. Modern sophisticated methods, in contrast, are robust to noise, but may be too slow for real-time processing of large images. This raises the following question, which is the focus of our paper: How well can one detect long edges in noisy images under severe computational constraints that allow only a fraction of all image pixels to be processed? We make several theoretical and practical contributions regarding this problem. We develop possibly the first sublinear algorithm to detect long straight edges in noisy images. In addition, we theoretically analyze the inevitable tradeoff between its detection performance and the allowed computational budget. Finally, we demonstrate its competitive performance on both simulated and real images. Inbal Horev, Boaz Nadler, Ery Arias-Castro, Meirav Galun, Ronen Basri |
SIAM J. Imaging Sci. | 5 |
| 2015 | A Global Approach for Solving Edge-Matching PuzzlesabstractWe consider apictorial edge-matching puzzles, in which the goal is to arrange a collection of puzzle pieces with colored edges so that the colors match along the edges of adjacent pieces. We devise an algebraic representation for this problem and provide conditions under which it exactly characterizes a puzzle. Using the new representation, we recast the combinatorial, discrete problem of solving puzzles as a global, polynomial system of equations with continuous variables. We further propose new algorithms for generating approximate solutions to the continuous problem by solving a sequence of convex relaxations. Shahar Z. Kovalsky, Daniel Glasner, Ronen Basri |
SIAM J. Imaging Sci. | 3 |
| 2015 | Stable Camera Motion Estimation Using Convex ProgrammingabstractWe study the inverse problem of estimating $n$ locations $\mathbf{t}_1, \mathbf{t}_2, \ldots, \mathbf{t}_n$ (up to global scale, translation, and negation) in $\mathbb{R}^d$ from noisy measurements of a subset of the (unsigned) pairwise lines that connect them, that is, from noisy measurements of $\pm \frac{\mathbf{t}_i - \mathbf{t}_j}{\|\mathbf{t}_i - \mathbf{t}_j \|_2}$ for some pairs $(i,j)$ (where the signs are unknown). This problem is at the core of the structure from motion (SfM) problem in computer vision, where the $\mathbf{t}_i$ represent camera locations in $\mathbb{R}^3$. The noiseless version of the problem, with exact line measurements, has been considered previously under the general title of parallel rigidity theory, mainly in order to characterize the conditions for unique realization of locations. For noisy pairwise line measurements, current methods tend to produce spurious solutions that are clustered around a few locations. This sensitivity of the location estimates is a well-known problem in SfM, especially for large, irregular collections of images. In this paper we introduce a semidefinite programming (SDP) formulation, specially tailored to overcome the clustering phenomenon. We further identify the implications of parallel rigidity theory for the location estimation problem to be well-posed, and prove exact (in the noiseless case) and stable location recovery results. We also formulate an alternating direction method to solve the resulting semidefinite program, and provide a distributed version of our formulation for large numbers of locations. Specifically for the camera location estimation problem, we formulate a pairwise line estimation method based on robust camera orientation and subspace estimation. Finally, we demonstrate the utility of our algorithm through experiments on real images. Onur Özyesil, Amit Singer, Ronen Basri |
SIAM J. Imaging Sci. | 3 |
| 2015 | Large-scale bounded distortion mappingsabstractWe propose an efficient algorithm for computing large-scale bounded distortion maps of triangular and tetrahedral meshes. Specifically, given an initial map, we compute a similar map whose differentials are orientation preserving and have bounded condition number. Inspired by alternating optimization and Gauss-Newton approaches, we devise a first order method which combines the advantages of both. On the one hand, its iterations are as computationally efficient as those of alternating optimization. On the other hand, it enjoys preferable convergence properties, associated with Gauss-Newton like approaches. We demonstrate the utility of the proposed approach in efficiently solving geometry processing problems, focusing on challenging large-scale problems. Shahar Z. Kovalsky, Noam Aigerman, Ronen Basri, Yaron Lipman |
ACM Trans. Graph. | 3 |
| 2014 | Controlling singular values with semidefinite programmingabstractControlling the singular values of n -dimensional matrices is often required in geometric algorithms in graphics and engineering. This paper introduces a convex framework for problems that involve singular values. Specifically, it enables the optimization of functionals and constraints expressed in terms of the extremal singular values of matrices. Towards this end, we introduce a family of convex sets of matrices whose singular values are bounded. These sets are formulated using Linear Matrix Inequalities (LMI), allowing optimization with standard convex Semidefinite Programming (SDP) solvers. We further show that these sets are optimal, in the sense that there exist no larger convex sets that bound singular values. A number of geometry processing problems are naturally described in terms of singular values. We employ the proposed framework to optimize and improve upon standard approaches. We experiment with this new framework in several applications: volumetric mesh deformations, extremal quasi-conformal mappings in three dimensions, non-rigid shape registration and averaging of rotations. We show that in all applications the proposed approach leads to algorithms that compare favorably to state-of-art algorithms. Shahar Z. Kovalsky, Noam Aigerman, Ronen Basri, Yaron Lipman |
ACM Trans. Graph. | 3 |
| 2014 | Feature Matching with Bounded DistortionabstractWe consider the problem of finding a geometrically consistent set of point matches between two images. We assume that local descriptors have provided a set of candidate matches, which may include many outliers. We then seek the largest subset of these correspondences that can be aligned perfectly using a nonrigid deformation that exerts a bounded distortion. We formulate this as a constrained optimization problem and solve it using a constrained, iterative reweighted least-squares algorithm. In each iteration of this algorithm we solve a convex quadratic program obtaining a globally optimal match over a subset of the bounded distortion transformations. We further prove that a sequence of such iterations converges monotonically to a critical point of our objective function. We show experimentally that this algorithm produces excellent results on a number of test sets, in comparison to several state-of-the-art approaches. Yaron Lipman, Stav Yagev, Roi Poranne, David Jacobs 0001, Ronen Basri |
ACM Trans. Graph. | 5 |
| 2012 | Viewpoint-aware object detection and continuous pose estimation
Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
Image Vis. Comput. | 4 |
| 2012 | Image Segmentation by Probabilistic Bottom-Up Aggregation and Cue IntegrationabstractWe present a bottom-up aggregation approach to image segmentation. Beginning with an image, we execute a sequence of steps in which pixels are gradually merged to produce larger and larger regions. In each step, we consider pairs of adjacent regions and provide a probability measure to assess whether or not they should be included in the same segment. Our probabilistic formulation takes into account intensity and texture distributions in a local area around each region. It further incorporates priors based on the geometry of the regions. Finally, posteriors based on intensity and texture cues are combined using “a mixture of experts” formulation. This probabilistic approach is integrated into a graph coarsening scheme, providing a complete hierarchical segmentation of the image. The algorithm complexity is linear in the number of the image pixels and it requires almost no user-tuned parameters. In addition, we provide a novel evaluation scheme for image segmentation algorithms, attempting to avoid human semantic considerations that are out of scope for segmentation algorithms. Using this novel evaluation scheme, we test our method and provide a comparison to several existing segmentation algorithms. Sharon Alpert, Meirav Galun, Achi Brandt, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Contour-based joint clustering of multiple segmentationsabstractWe present an unsupervised, shape-based method for joint clustering of multiple image segmentations. Given two or more closely-related images, such as nearby frames in a video sequence or images of the same scene taken under different lighting conditions, our method generates a joint segmentation of the images. We introduce a novel contour-based representation that allows us to cast the shape-based joint clustering problem as a quadratic semi-assignment problem. Our score function is additive. We use complex-valued affinities to assess the quality of matching the edge elements at the exterior bounding contour of clusters, while ignoring the contributions of elements that fall in the interior of the clusters. We further combine this contour-based score with region information and use a linear programming relaxation to solve for the joint clusters. We evaluate our approach on the occlusion boundary data-set of Stein et al. Daniel Glasner, Shiv Vitaladevuni, Ronen Basri |
CVPR | 3 |
| 2011 | Viewpoint-aware object detection and pose estimationabstractWe describe an approach to category-level detection and viewpoint estimation for rigid 3D objects from single 2D images. In contrast to many existing methods, we directly integrate 3D reasoning with an appearance-based voting architecture. Our method relies on a nonparametric representation of a joint distribution of shape and appearance of the object class. Our voting method employs a novel parametrization of joint detection and viewpoint hypothesis space, allowing efficient accumulation of evidence. We combine this with a re-scoring and refinement mechanism, using an ensemble of view-specific Support Vector Machines. We evaluate the performance of our approach in detection and pose estimation of cars on a number of benchmark datasets. Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
ICCV | 4 |
| 2011 | Approximate Nearest Subspace SearchabstractSubspaces offer convenient means of representing information in many pattern recognition, machine vision, and statistical learning applications. Contrary to the growing popularity of subspace representations, the problem of efficiently searching through large subspace databases has received little attention in the past. In this paper, we present a general solution to the problem of Approximate Nearest Subspace search. Our solution uniformly handles cases where the queries are points or subspaces, where query and database elements differ in dimensionality, and where the database contains subspaces of different dimensions. To this end, we present a simple mapping from subspaces to points, thus reducing the problem to the well-studied Approximate Nearest Neighbor problem on points. We provide theoretical proofs of correctness and error bounds of our construction and demonstrate its capabilities on synthetic and real data. Our experiments indicate that an approximate nearest subspace can be located significantly faster than the nearest subspace, with little loss of accuracy. Ronen Basri, Tal Hassner, Lihi Zelnik-Manor |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | 3D Face Reconstruction from a Single Image Using a Single Reference Face ShapeabstractHuman faces are remarkably similar in global properties, including size, aspect ratio, and location of main features, but can vary considerably in details across individuals, gender, race, or due to facial expression. We propose a novel method for 3D shape recovery of faces that exploits the similarity of faces. Our method obtains as input a single image and uses a mere single 3D reference model of a different person's face. Classical reconstruction methods from single images, i.e., shape-from-shading, require knowledge of the reflectance properties and lighting as well as depth values for boundary conditions. Recent methods circumvent these requirements by representing input faces as combinations (of hundreds) of stored 3D models. We propose instead to use the input image as a guide to "mold" a single reference model to reach a reconstruction of the sought 3D shape. Our method assumes Lambertian reflectance and uses harmonic representations of lighting. It has been tested on images taken under controlled viewing conditions as well as on uncontrolled images downloaded from the Internet, demonstrating its accuracy and robustness under a variety of imaging conditions and overcoming significant differences in shape between the input and reference individuals including differences in facial expressions, gender, and race. Ira Kemelmacher-Shlizerman, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Co-clustering of image segments using convex optimization applied to EM neuronal reconstructionabstractThis paper addresses the problem of jointly clustering two segmentations of closely correlated images. We focus in particular on the application of reconstructing neuronal structures in over-segmented electron microscopy images. We formulate the problem of co-clustering as a quadratic semi-assignment problem and investigate convex relaxations using semidefinite and linear programming. We further introduce a linear programming method with manageable number of constraints and present an approach for learning the cost function. Our method increases computational efficiency by orders of magnitude while maintaining accuracy, automatically finds the optimal number of clusters, and empirically tends to produce binary assignment solutions. We illustrate our approach in simulations and in experiments with real EM data. Shiv Vitaladevuni, Ronen Basri |
CVPR | 2 |
| 2010 | Detecting Faint Curved Edges in Noisy Images
Sharon Alpert, Meirav Galun, Boaz Nadler, Ronen Basri |
ECCV (4) | 4 |
| 2009 | Visibility constraints on features of 3D objectsabstractTo recognize three-dimensional objects it is important to model how their appearances can change due to changes in viewpoint. A key aspect of this involves understanding which object features can be simultaneously visible under different viewpoints. We address this problem in an image-based framework, in which we use a limited number of images of an object taken from unknown viewpoints to determine which subsets of features might be simultaneously visible in other views. This leads to the problem of determining whether a set of images, each containing a set of features, is consistent with a single 3D object. We assume that each feature is visible from a disk of viewpoints on the viewing sphere. In this case we show the problem is NP-hard in general, but can be solved efficiently when all views come from a circle on the viewing sphere. We also give iterative algorithms that can handle noisy data and converge to locally optimal solutions in the general case. Our techniques can also be used to recover viewpoint information from the set of features that are visible in different images. We show that these algorithms perform well both on synthetic data and images from the COIL dataset. Ronen Basri, Pedro F. Felzenszwalb, Ross B. Girshick, David Jacobs 0001, Caroline J. Klivans |
CVPR | 1 |
| 2009 | Constructing implicit 3D shape models for pose estimationabstractWe present a system that constructs “implicit shape models” for classes of rigid 3D objects and utilizes these models to estimating the pose of class instances in single 2D images. We use the framework of implicit shape models to construct a voting procedure that allows for 3D transformations and projection and accounts for self occlusion. The model is comprised of a collection of learned features, their 3D locations, their appearances in different views, and the set of views in which they are visible. We further learn the parameters of a model from training images by applying a method that relies on factorization. We demonstrate the utility of the constructed models by applying them in pose estimation experiments to recover the viewpoint of class instances. Mica Arie-Nachimson, Ronen Basri |
ICCV | 2 |
| 2009 | Shape Based Detection and Top-Down Delineation Using Image Segments
Lena Gorelick, Ronen Basri |
Int. J. Comput. Vis. | 2 |
| 2008 | A two-frame theory of motion, lighting and shapeabstractThis paper explores how shape, motion, and lighting interact in the case of a two-frame motion sequence. We consider a rigid object with Lambertian reflectance properties undergoing small motion with respect to both a camera and a stationary point light source. Assuming orthographic projection, we derive a single, first order quasilinear partial differential equation that relates shape, motion, and lighting, while eliminating out the albedo. We show how this equation can be solved, when the motion and lighting parameters are known, to produce a 3D reconstruction of the object. A solution is obtained using the method of characteristics and can be refined by adding regularization. We further show that both smooth bounding contours as well as surface markings can be used to derive Dirichlet boundary conditions. Experimental results demonstrate the quality of this reconstruction. Ronen Basri, Darya Frolova |
CVPR | 1 |
| 2008 | 3D shape reconstruction of Mooney facesabstractTwo-tone (ldquoMooneyrdquo) images seem to arouse vivid 3D percept of faces, both familiar and unfamiliar, despite their seemingly poor content. Recent psychological and fMRI studies suggest that this percept is guided primarily by top-down procedures in which recognition precedes reconstruction. In this paper we investigate this hypothesis from a mathematical standpoint. We show that indeed, under standard shape from shading assumptions, a Mooney image can give rise to multiple different 3D reconstructions even if reconstruction is restricted to the Mooney transition curve (the boundary curve between black and white) alone. We then use top-down reconstruction methods to recover the shape of novel faces from single Mooney images exploiting prior knowledge of the structure of at least one face of a different individual. We apply these methods to thresholded images of real faces and compare the reconstruction quality relative to reconstruction from gray level images. Ira Kemelmacher-Shlizerman, Ronen Basri, Boaz Nadler |
CVPR | 2 |
| 2007 | Image Segmentation by Probabilistic Bottom-Up Aggregation and Cue IntegrationabstractWe present a parameter free approach that utilizes multiple cues for image segmentation. Beginning with an image, we execute a sequence of bottom-up aggregation steps in which pixels are gradually merged to produce larger and larger regions. In each step we consider pairs of adjacent regions and provide a probability measure to assess whether or not they should be included in the same segment. Our probabilistic formulation takes into account intensity and texture distributions in a local area around each region. It further incorporates priors based on the geometry of the regions. Finally, posteriors based on intensity and texture cues are combined using a mixture of experts formulation. This probabilistic approach is integrated into a graph coarsening scheme providing a complete hierarchical segmentation of the image. The algorithm complexity is linear in the number of the image pixels and it requires almost no user-tuned parameters. We test our method on a variety of gray scale images and compare our results to several existing segmentation algorithms. Sharon Alpert, Meirav Galun, Ronen Basri, Achi Brandt |
CVPR | 3 |
| 2007 | Approximate Nearest Subspace Search with Applications to Pattern RecognitionabstractLinear and affine subspaces are commonly used to describe appearance of objects under different lighting, viewpoint, articulation, and identity. A natural problem arising from their use is - given a query image portion represented as a point in some high dimensional space - find a subspace near to the query. This paper presents an efficient solution to the approximate nearest subspace problem for both linear and affine subspaces. Our method is based on a simple reduction to the problem of nearest point search, and can thus employ tree based search or locality sensitive hashing to find a near subspace. Further speedup may be achieved by using random projections to lower the dimensionality of the problem. We provide theoretical proofs of correctness and error bounds of our construction and demonstrate its capabilities on synthetic and real data. Our experiments demonstrate that an approximate nearest subspace can be located significantly faster than the exact nearest subspace, while at the same time it can find better matches compared to a similar search on points, in the presence of variations due to viewpoint, lighting etc. Ronen Basri, Tal Hassner, Lihi Zelnik-Manor |
CVPR | 1 |
| 2007 | Multiscale Edge Detection and Fiber Enhancement Using Differences of Oriented MeansabstractWe present an algorithm for edge detection suitable for both natural as well as noisy images. Our method is based on efficient multiscale utilization of elongated filters measuring the difference of oriented means of various lengths and orientations, along with a theoretical estimation of the effect of noise on the response of such filters. We use a scale adaptive threshold along with a recursive decision process to reveal the significant edges of all lengths and orientations and to localize them accurately even in low-contrast and very noisy images. We further use this algorithm for fiber detection and enhancement by utilizing stochastic completion-like process from both sides of a fiber. Our algorithm relies on an efficient multiscale algorithm for computing all "significantly different" oriented means in an image in O(N log rho), where N is the number of pixels, and p is the length of the longest structure of interest. Experimental results on both natural and noisy images are presented. Meirav Galun, Ronen Basri, Achi Brandt |
ICCV | 2 |
| 2007 | Prior Knowledge Driven Multiscale Segmentation of Brain MRI
Ayelet Akselrod-Ballin, Meirav Galun, Moshe John Gomori, Achi Brandt, Ronen Basri |
MICCAI (2) | 5 |
| 2007 | Rediscovering secondary structures as network motifs - an unsupervised learning approachabstractMOTIVATION: Secondary structures are key descriptors of a protein fold and its topology. In recent years, they facilitated intensive computational tasks for finding structural homologues, fold prediction and protein design. Their popularity stems from an appealing regularity in patterns of geometry and chemistry. However, the definition of secondary structures is of subjective nature. An unsupervised de-novo discovery of these structures would shed light on their nature, and improve the way we use these structures in algorithms of structural bioinformatics. METHODS: We developed a new method for unsupervised partitioning of undirected graphs, based on patterns of small recurring network motifs. Our input was the network of all H-bonds and covalent interactions of protein backbones. This method can be also used for other biological and non-biological networks. RESULTS: In a fully unsupervised manner, and without assuming any explicit prior knowledge, we were able to rediscover the existence of conventional alpha-helices, parallel beta-sheets, anti-parallel sheets and loops, as well as various non-conventional hybrid structures. The relation between connectivity and crystallographic temperature factors establishes the existence of novel secondary structures. Barak Raveh, Ofer Rahat, Ronen Basri, Gideon Schreiber |
Bioinform. | 3 |
| 2007 | Photometric Stereo with General, Unknown Lighting
Ronen Basri, David Jacobs 0001, Ira Kemelmacher-Shlizerman |
Int. J. Comput. Vis. | 1 |
| 2007 | Actions as Space-Time ShapesabstractHuman action in video sequences can be seen as silhouettes of a moving torso and protruding limbs undergoing articulated motion. We regard human actions as three-dimensional shapes induced by the silhouettes in the space-time volume. We adopt a recent approach for analyzing 2D shapes and generalize it to deal with volumetric space-time action shapes. Our method utilizes properties of the solution to the Poisson equation to extract space-time features such as local space-time saliency, action dynamics, shape structure and orientation. We show that these features are useful for action recognition, detection and clustering. The method is fast, does not require video alignment and is applicable in (but not limited to) many scenarios where the background is known. Moreover, we demonstrate the robustness of our method to partial occlusions, non-rigid deformations, significant changes in scale and viewpoint, high irregularities in the performance of an action, and low quality video. Lena Gorelick, Moshe Blank, Eli Shechtman, Michal Irani, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2007 | Direct visibility of point setsabstractThis paper proposes a simple and fast operator, the "Hidden" Point Removal operator, which determines the visible points in a point cloud, as viewed from a given viewpoint. Visibility is determined without reconstructing a surface or estimating normals. It is shown that extracting the points that reside on the convex hull of a transformed point cloud, amounts to determining the visible points. This operator is general - it can be applied to point clouds at various dimensions, on both sparse and dense point clouds, and on viewpoints internal as well as external to the cloud. It is demonstrated that the operator is useful in visualizing point clouds, in view-dependent reconstruction and in shadow casting. Sagi Katz, Ayellet Tal, Ronen Basri |
ACM Trans. Graph. | 3 |
| 2006 | An Integrated Segmentation and Classification Approach Applied to Multiple Sclerosis AnalysisabstractWe present a novel multiscale approach that combines segmentation with classification to detect abnormal brain structures in medical imagery, and demonstrate its utility in detecting multiple sclerosis lesions in 3D MRI data. Our method uses segmentation to obtain a hierarchical decomposition of a multi-channel, anisotropic MRI scan. It then produces a rich set of features describing the segments in terms of intensity, shape, location, and neighborhood relations. These features are then fed into a decision tree-based classifier, trained with data labeled by experts, enabling the detection of lesions in all scales. Unlike common approaches that use voxel-by-voxel analysis, our system can utilize regional properties that are often important for characterizing abnormal brain structures. We provide experiments showing successful detections of lesions in both simulated and real MR images. Ayelet Akselrod-Ballin, Meirav Galun, Ronen Basri, Achi Brandt, Moshe John Gomori, Massimo Filippi, Paola Valsasina |
CVPR (1) | 3 |
| 2006 | Molding Face Shapes by Example
Ira Kemelmacher-Shlizerman, Ronen Basri |
ECCV (1) | 2 |
| 2006 | Atlas Guided Identification of Brain Structures by Combining 3D Segmentation and SVM Classification
Ayelet Akselrod-Ballin, Meirav Galun, Moshe John Gomori, Ronen Basri, Achi Brandt |
MICCAI (2) | 4 |
| 2006 | Shape Representation and Classification Using the Poisson EquationabstractWe present a novel approach that allows us to reliably compute many useful properties of a silhouette. Our approach assigns, for every internal point of the silhouette, a value reflecting the mean time required for a random walk beginning at the point to hit the boundaries. This function can be computed by solving Poisson's equation, with the silhouette contours providing boundary conditions. We show how this function can be used to reliably extract various shape properties including part structure and rough skeleton, local orientation and aspect ratio of different parts, and convex and concave sections of the boundaries. In addition to this, we discuss properties of the solution and show how to efficiently compute this solution using multigrid algorithms. We demonstrate the utility of the extracted properties by using them for shape classification and retrieval. Lena Gorelick, Meirav Galun, Eitan Sharon, Ronen Basri, Achi Brandt |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2005 | Multiscale Segmentation by Combining Motion and Intensity CuesabstractWe present a multiscale method for motion segmentation. Our method begins with local, ambiguous optical flow measurements. It uses a process of aggregation to resolve the ambiguities and reach reliable estimates of the motion. In addition, as the aggregation process proceeds and larger aggregates are identified it employs a progressively more complex model to describe the motion. In particular, we proceed by recovering translational motion at fine levels, through affine transformation at intermediate levels, to 3D motion (described by a fundamental matrix) at the coarsest levels. Finally, the method is integrated with a segmentation method that uses intensity cues. We further demonstrate the utility of the method on both random dot and real motion sequences. Meirav Galun, Alexander Apartsin, Ronen Basri |
CVPR (1) | 3 |
| 2005 | Indexing with Unknown Illumination and PoseabstractThe task of identifying 3D objects in 2D images is difficult due to variation in objects' appearance with changes in pose and lighting. The task is further complicated by the presence of occlusion and clutter. Shape indexing is a method for rapid association between features identified in an image and their corresponding 3D features stored in a database. Previous indexing methods ignored variations due to lighting, restricting the approach to polyhedral objects. In this paper, we further develop these methods to handle variations in both pose and lighting. We focus on rigid objects undergoing a scaled-orthographic projection and use spherical harmonics to represent lighting. The resulting integrated algorithm can recognize 3D objects from a single input image; furthermore, it recovers the pose and lighting of each familiar object in the given image. The algorithm has been tested on a database of real objects, demonstrating its performance on cluttered scenes under a variety of poses and illumination conditions. Ira Kemelmacher-Shlizerman, Ronen Basri |
CVPR (1) | 2 |
| 2005 | Actions as Space-Time ShapesabstractHuman action in video sequences can be seen as silhouettes of a moving torso and protruding limbs undergoing articulated motion. We regard human actions as three-dimensional shapes induced by the silhouettes in the space-time volume. We adopt a recent approach by Gorelick et al. (2004) for analyzing 2D shapes and generalize it to deal with volumetric space-time action shapes. Our method utilizes properties of the solution to the Poisson equation to extract space-time features such as local space-time saliency, action dynamics, shape structure and orientation. We show that these features are useful for action recognition, detection and clustering. The method is fast, does not require video alignment and is applicable in (but not limited to) many scenarios where the background is known. Moreover, we demonstrate the robustness of our method to partial occlusions, non-rigid deformations, significant changes in scale and viewpoint, high irregularities in the performance of an action and low quality video Moshe Blank, Lena Gorelick, Eli Shechtman, Michal Irani, Ronen Basri |
ICCV | 5 |
| 2005 | Minimal-Cut Model CompositionabstractConstructing new, complex models is often done by reusing parts of existing models, typically by applying a sequence of segmentation, alignment and composition operations. Segmentation, either manual or automatic, is rarely adequate for this task, since it is applied to each model independently, leaving it to the user to trim the models and determine where to connect them. In this paper we propose a new composition tool. Our tool obtains as input two models, aligned either manually or automatically, and a small set of constraints indicating which portions of the two models should be preserved in the final output. It then automatically negotiates the best location to connect the models, trimming and stitching them as required to produce a seamless result. We offer a method based on the graph theoretic minimal cut as a means of implementing this new tool. We describe a system intended for both expert and novice users, allowing easy and flexible control over the composition result. In addition, we show our method to be well suited for a variety of model processing applications such as model repair, hole filling, and piecewise rigid deformations. Tal Hassner, Lihi Zelnik-Manor, George Leifman, Ronen Basri |
SMI | 4 |
| 2004 | Shape Representation and Classification Using the Poisson Equation
Lena Gorelick, Meirav Galun, Eitan Sharon, Ronen Basri, Achi Brandt |
CVPR (2) | 4 |
| 2004 | Accuracy of Spherical Harmonic Approximations for Images of Lambertian Objects under Far and Near Lighting
Darya Frolova, Denis Simakov, Ronen Basri |
ECCV (1) | 3 |
| 2003 | Texture Segmentation by Multiscale Aggregation of Filter Responses and Shape ElementsabstractTexture segmentation is a difficult problem, as is apparent from camouflage pictures. A textured region can contain texture elements of various sizes, each of which can itself be textured. We approach this problem using a bottom-up aggregation framework that combines structural characteristics of texture elements with filter responses. Our process adaptively identifies the shape of texture elements and characterize them by their size, aspect ratio, orientation, brightness, etc., and then uses various statistics of these properties to distinguish between different textures. At the same time our process uses the statistics of filter responses to characterize textures. In our process the shape measures and the filter responses crosstalk extensively. In addition, a top-down cleaning process is applied to avoid mixing the statistics of neighboring segments. We tested our algorithm on real images and demonstrate that it can accurately segment regions that contain challenging textures. Meirav Galun, Eitan Sharon, Ronen Basri, Achi Brandt |
ICCV | 3 |
| 2003 | Dense Shape Reconstruction of a Moving Object under Arbitrary, Unknown LightingabstractWe present a method for shape reconstruction from several images of a moving object. The reconstruction is dense (up to image resolution). The method assumes that the motion is known, e.g., by tracking a small number of feature points on the object. The object is assumed Lambertian (completely matte), light sources should not be very close to the object but otherwise arbitrary, and no knowledge of lighting conditions is required. An object changes its appearance significantly when it changes its orientation relative to light sources, causing violation of the common brightness constancy assumption. While a lot of effort is devoted to deal with this violation, we demonstrate how to exploit it to recover 3D structure from 2D images. We propose a new correspondence measure that enables point matching across views of a moving object. The method has been tested both on computer simulated examples and on a real object. Denis Simakov, Darya Frolova, Ronen Basri |
ICCV | 3 |
| 2003 | Lambertian Reflectance and Linear SubspacesabstractWe prove that the set of all Lambertian reflectance functions (the mapping from surface normals to intensities) obtained with arbitrary distant light sources lies close to a 9D linear subspace. This implies that, in general, the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace, explaining prior empirical results. We also provide a simple analytic characterization of this linear space. We obtain these results by representing lighting using spherical harmonics and describing the effects of Lambertian materials as the analog of a convolution. These results allow us to construct algorithms for object recognition based on linear methods as well as algorithms that use convex optimization to enforce nonnegative lighting functions. We also show a simple way to enforce nonnegative lighting when the images of an object lie near a 4D linear space. We apply these algorithms to perform face recognition by finding the 3D model that best matches a 2D query image. Ronen Basri, David Jacobs 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Inferring region salience from binary and gray-level images
Yossi Cohen, Ronen Basri |
Pattern Recognit. | 2 |
| 2001 | Photometric Stereo with General, Unknown LightingabstractWork on photometric stereo has shown how to recover the shape and reflectance properties of an object using multiple images taken with a fixed viewpoint and variable lighting conditions. This work has primarily relied on the presence of a single point source of light in each image. The authors show how to perform photometric stereo, assuming that all lights in a scene are isotropic and distant from the object but otherwise unconstrained. Lighting in each image may be an unknown and arbitrary combination of diffuse, point and extended sources. Our work is based on recent results showing that for Lambertian objects, general lighting conditions can be represented using low order spherical harmonics. Using this representation, we can recover shape by performing a simple optimization in a low-dimensional space. We also analyze the shape ambiguities that arise in such a representation. Ronen Basri, David Jacobs 0001 |
CVPR (2) | 1 |
| 2001 | Segmentation and Boundary Detection Using Multiscale Intensity MeasurementsabstractImage segmentation is difficult because objects may differ from their background by any of a variety of properties that can be observed in some, but often not all scales. A further complication is that coarse measurements, applied to the image for detecting these properties, often average over properties of neighboring segments, making it difficult to separate the segments and to reliably detect their boundaries. Below we present a method for segmentation that generates and combines multiscale measurements of intensity contrast, texture differences, and boundary integrity. The method is based on our former algorithm SWA, which efficiently detects segments that optimize a normalized-cut like measure by recursively coarsening a graph reflecting similarities between intensities of neighboring pixels. In this process aggregates of pixels of increasing size are gradually collected to form segments. We intervene in this process by computing properties of the aggregates and modifying the graph to reflect these coarse scale measurements. This allows us to detect regions that differ by fine as well as coarse properties, and to accurately locate their boundaries. Furthermore, by combining intensity differences with measures of boundary integrity across neighboring aggregates we can detect regions separated by weak, yet consistent edges. Eitan Sharon, Achi Brandt, Ronen Basri |
CVPR (1) | 3 |
| 2001 | Lambertian Reflectance and Linear SubspacesabstractWe prove that the set of all reflectance functions (the mapping from surface normals to intensities) produced by Lambertian objects under distant, isotropic lighting lies close to a 9D linear subspace. This implies that the images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately with a low-dimensional linear subspace, explaining prior empirical results. We also provide a simple analytic characterization of this linear space. We obtain these results by representing lighting using spherical harmonics and describing the effects of Lambertian materials as the analog of a convolution. These results allow us to construct algorithms for object recognition based on linear methods as well as algorithms that use convex optimization to enforce non-negative lighting functions. Ronen Basri, David Jacobs 0001 |
ICCV | 1 |
| 2001 | Projective Alignment with RegionsabstractWe have previously proposed (Basri and Jacobs, 1999, and Jacobs and Basri, 1999) an approach to recognition that uses regions to determine the pose of objects while allowing for partial occlusion of the regions. Regions introduce an attractive alternative to existing global and local approaches, since, unlike global features, they can handle occlusion and segmentation errors, and unlike local features they are not as sensitive to sensor errors, and they are easier to match. The region-based approach also uses image information directly, without the construction of intermediate representations, such as algebraic descriptions, which may be difficult to reliably compute. We further analyze properties of the method for planar objects undergoing projective transformations. In particular, we prove that three visible regions are sufficient to determine the transformation uniquely and that for a large class of objects, two regions are insufficient for this purpose. However, we show that when several regions are available, the pose of the object can generally be recovered even when some or all regions are significantly occluded. Our analysis is based on investigating the flow patterns of points under projective transformations in the presence of fixed points. Ronen Basri, David Jacobs 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Fast Multiscale Image SegmentationabstractWe introduce a fast, multiscale algorithm for image segmentation. Our algorithm uses modern numeric techniques to find an approximate solution to normalized cut measures in time that is linear in the size of the image with only a few dozen operations per pixel. In just one pass the algorithm provides a complete hierarchical decomposition of the image into segments. The algorithm detects the segments by applying a process of recursive coarsening in which the same minimization problem is represented with fewer and fewer variables producing an irregular pyramid. During this coarsening process we may compute additional internal statistics of the emerging segments and use these statistics to facilitate the segmentation process. Once the pyramid is completed it is scanned from the top down to associate pixels close to the boundaries of segments with the appropriate segment. The algorithm is inspired by algebraic multigrid (AMG) solvers of minimization problems of heat or electric networks. We demonstrate the algorithm by applying it to real images. Eitan Sharon, Achi Brandt, Ronen Basri |
CVPR | 3 |
| 2000 | Separation of Transparent Layers using Focus
Yoav Y. Schechner, Nahum Kiryati, Ronen Basri |
Int. J. Comput. Vis. | 3 |
| 2000 | Completion Energies and ScaleabstractThe detection of smooth curves in images and their completion over gaps are two important problems in perceptual grouping. We examine the notion of completion energy of curve elements, showing, and exploiting its intrinsic dependence on length and width scales. We introduce a fast method for computing the most likely completion between two elements, by developing novel analytic approximations and a fast numerical procedure for computing the curve of least energy. We then use our newly developed energies to find the most likely completions in images through a generalized summation of induction fields. This is done through multiscale procedures, i.e., separate processing at different scales with some interscale interactions. Such procedures allow the summation of all induction fields to be done in a total of only O(N log N) operations, where N is the number of pixels in the image. More important, such procedures yield a more realistic dependence of the induction field on the length and width scales: the field of a long element is very different from the sum of the fields of its composing short segments. Eitan Sharon, Achi Brandt, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1999 | Projective Alignment with RegionsabstractWe consider a recent approach to recognition that uses regions to determine the pose of objects while allowing for partial occlusion of the regions. We further analyze properties of the method for planar objects undergoing projective transformations. We prove that three visible regions are sufficient to determine the transformation uniquely, and that for a large class of objects two regions are insufficient. However, we show that when several regions are available, the pose of the object can generally be recovered even when all but two regions are significantly occluded. Our analysis is based on investigating the flow patterns of points under projective transformations in the presence of fixed points. Ronen Basri, David Jacobs 0001 |
ICCV | 1 |
| 1999 | Image-Based Robot Navigation Under the Perspective ModelabstractIn a previous paper (1998) we presented a method for image-based navigation by which a robot can navigate to desired positions and orientations in 3D space specified by single images taken from these positions. In this paper we further investigate the method and develop robust algorithms for navigation assuming the perspective projection model. In particular, we develop a tracking algorithm that exploits our knowledge of the motion performed by the robot at every step. This algorithm allows us to maintain correspondences between frames and eliminate false correspondences. We combine this tracking algorithm with an iterative optimization procedure to accurately recover the displacement of the robot from the target. Our method for navigation is attractive since it does not require a 3D model of the environment. We demonstrate the robustness of our method by applying it to a six degree of freedom robot arm. Ronen Basri, Ehud Rivlin, Ilan Shimshoni |
ICRA | 1 |
| 1999 | When is it Possible to Identify 3D Objects From Single Images Using Class Constraints?
Ronen Basri, Yael Moses |
Int. J. Comput. Vis. | 1 |
| 1999 | Visual Homing: Surfing on the Epipoles
Ronen Basri, Ehud Rivlin, Ilan Shimshoni |
Int. J. Comput. Vis. | 1 |
| 1999 | 3-D to 2-D Pose Determination with Regions
David Jacobs 0001, Ronen Basri |
Int. J. Comput. Vis. | 2 |
| 1999 | A Geometric Interpretation of Weak-Perspective MotionabstractWe present a geometric interpretation of the problem of motion recovery from three weak-perspective images. Our interpretation is based on reducing the problem of estimating the motion to a problem of finding triangles on a sphere whose angles are known. Using this geometric interpretation, a simple method to completely recover the motion parameters using three images is developed. The results of running the algorithm on real images are presented. In addition, we describe which of the various motion parameters can be recovered already from two images. Ilan Shimshoni, Ronen Basri, Ehud Rivlin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Clustering Appearances of 3D ObjectsabstractWe introduce a method for unsupervised clustering of images of 3D objects. Our method examines the space of all images and partitions the images into sets that form smooth and parallel surfaces in this space. It further uses sequences of images to obtain more reliable clustering. Finally, since our method relies on a non-Euclidean similarity measure we introduce algebraic techniques for estimating local properties of these surfaces without first embedding the images in a Euclidean space. We demonstrate our method by applying it to a large database of images. Ronen Basri, Dan Roth 0001, David Jacobs 0001 |
CVPR | 1 |
| 1998 | Comparing Images under Variable IlluminationabstractWe consider the problem of determining whether two images come from different objects or the same object in the same pose, but under different illumination conditions. We show that this problem cannot be solved using hard constraints: even using a Lambertian reflectance model, there is always an object and a pair of lighting conditions consistent with any two images. Nevertheless, we show that for point sources and objects with Lambertian reflectance, the ratio of two images from the same object is simpler than the ratio of images from different objects. We also show that the ratio of the two images provides two of the three distinct values in the Hessian matrix of the object's surface. Using these observations, we develop a simple measure for matching images under variable illumination, comparing its performance to other existing methods on a database of 450 images of 10 individuals. David Jacobs 0001, Peter N. Belhumeur, Ronen Basri |
CVPR | 3 |
| 1998 | When is it Possible to Identify 3D Objects from Single Images Using Class Constraints?abstractOne approach to recognizing objects seen from arbitrary viewpoint is by extracting invariant properties of the objects from single images. Such properties are found in images of 3D objects only when the objects are constrained to belong to certain classes (e.g., bilaterally symmetric objects). Existing studies that follow this approach propose how to compute invariant representations for a handful of classes of objects. A fundamental question regarding the invariance approach is whether it call be applied to a wide range of classes. To answer this question it is essential to study the set of classes for which invariance exists. This paper introduces a new method for determining the existence of invariance for classes of objects together with the set of images from which these invariance can be computed. We develop algebraic tests that, given a class of objects undergoing affine projection, determine whether the objects in the class can be identified from single images. In addition, these tests allow us to determine the sell of views of the objects which are degenerate. We apply these tests to several classes of objects and determine which of them is identifiable and which of their views are degenerate. Ronen Basri, Yael Moses |
ICCV | 1 |
| 1998 | Visual Homing: Surfing on the EpipolesabstractWe introduce a novel method for visual homing. Using this method a robot can be sent to desired positions and orientations in 3-D space specified by single images taken from these positions. Our method determines the path of the robot on-line. The starting position of the robot is not constrained, and a 3-D model of the environment is not required. The method is based on recovering the epipolar geometry relating the current image taken by the robot and the target image. Using the epipolar geometry, most of the parameters which specify the differences in position and orientation of the camera between the two images are recovered. However, since not all of the parameters can be recovered from two images, we have developed specific methods to bypass these missing parameters and resolve the ambiguities that exist. We present two homing algorithms for two standard projection models, weak and full perspective. We have performed simulations and real experiments which demonstrate the robustness of the method and that the algorithms always converge to the target pose. Ronen Basri, Ehud Rivlin, Ilan Shimshoni |
ICCV | 1 |
| 1998 | Separation of Transparent Layers Using FocusabstractConsider situations where the depth at each point in the scene is multi-valued due to the presence of a virtual image semi-reflected by a transparent surface. The semi-reflected image is linearly superimposed on the image of the object that is behind the transparent surface. A novel approach is proposed for the recovery of the superimposed layers. By searching for the images in which either of the objects (layers) is focused, the transparent areas are detected and an estimate of the depth map of each layer is obtained. As a result of the focusing, an initial separation of the layers is achieved. The separation is enhanced via mutual blurring of the perturbing components in the images, based on the depths estimate and the parameters of the imaging system. Yoav Y. Schechner, Nahum Kiryati, Ronen Basri |
ICCV | 3 |
| 1998 | Extracting Salient Curves from Images: An Analysis of the Saliency Network
Tao Daniel Alter, Ronen Basri |
Int. J. Comput. Vis. | 2 |
| 1998 | Efficient determination of shape from multiple images containing partial information
Ronen Basri, Adam J. Grove, David Jacobs 0001 |
Pattern Recognit. | 1 |
| 1997 | 3-D to 2-D Recognition with RegionsabstractThis paper presents a novel approach to parts-based object recognition in the presence of occlusion. We focus on the problem of determining the pose of a 3-D object from a single 2-D image when convex parts of the object have been matched to corresponding regions in the image. We consider three types of occlusions: self-occlusion, occlusions whose locus is identified in the image, and completely arbitrary occlusions. We derive efficient algorithms for the first two cases, and characterize their performance. For the last case, we prove that the problem of finding valid poses is computationally hard, but provide an efficient, approximate algorithm. This work generalizes our previous work on region-based object recognition, which focused on the case of planar models. David Jacobs 0001, Ronen Basri |
CVPR | 2 |
| 1997 | Completion Energies and ScaleabstractThe detection of smooth curves in images and their completion over gaps are two important problems in perceptual grouping. In this paper we examine the notion of completion energy and introduce a fast method to compute the most likely completions in images. Specifically we develop two novel analytic approximations to the curve of least energy. In addition, we introduce a fast numerical method to compute the curve of least energy, and show that our approximations are obtained at early stages of this numerical computation. We then use our newly developed energies to find the most likely completions in images through a generalized summation of induction fields. Since in practice edge elements are obtained by applying filters of certain widths and lengths to the image, we adjust our computation to take these parameters into account. Finally, we show that, due to the smoothness of the kernel of summation, the process of summing induction fields can be run in time that is linear in the number of different edge elements in the image, or in O(N log N) where N is the number of pixels in the image, using multigrid methods. Eitan Sharon, Achi Brandt, Ronen Basri |
CVPR | 3 |
| 1997 | Constancy and Similarity
Ronen Basri, David Jacobs 0001 |
Comput. Vis. Image Underst. | 1 |
| 1997 | Recognition Using Region Correspondences
Ronen Basri, David Jacobs 0001 |
Int. J. Comput. Vis. | 1 |
| 1996 | Extracting Salient Curves from Images: An Analysis of the Saliency NetworkabstractThe Saliency Network proposed by Shashua and Ullman (1988) is a well-known approach to the problem of extracting salient curves from images while performing gap completion. This paper analyzes the Saliency Network. Although the network is attractive for a number reasons, our analysis reveals certain weaknesses with the method. In particular, we show cases in which the most salient element does not lie on the perceptually most salient curve. Furthermore, the saliency measure may change its preferences when curves are scaled uniformly. Also, for certain fragmented curves the measure prefers large gaps over a few small gaps of the same total size. We analyze the time complexity required by the method and discuss problems due to coarse sampling of the range of possible orientations. We show that with proper sampling the complexity of the network becomes cubic in the size of the network. Finally, we consider the possibility of using the Saliency Network for grouping. We show that the Saliency Network recovers the most salient curve efficiently, but it has problems with identifying any salient curve other than the most salient one. Tao Daniel Alter, Ronen Basri |
CVPR | 2 |
| 1996 | Efficient determination of shape from multiple images containing partial informationabstractWe consider the problem of reconstructing the shape of an object from multiple images related by translations, when only small portions of the object can be observed in each image. Lindenbaum and Bruckstein (1988) have considered this problem in the specific case where the translating object is seen by small sensors, for application to the understanding of insect vision. Their solution is limited by the fact that its run time is exponential in the number of images and sensors. We show that the problem can be solved in time that is polynomial in the number of sensors, but is in fact NP complete when the number of images is unbounded. We therefore consider the special case of convex objects, which we can solve efficiently even when many images are used. Ronen Basri, Adam J. Grove, David Jacobs 0001 |
ICPR | 1 |
| 1996 | Recognition by prototypes
Ronen Basri |
Int. J. Comput. Vis. | 1 |
| 1996 | Paraperspective = affine
Ronen Basri |
Int. J. Comput. Vis. | 1 |
| 1996 | Distance Metric Between 3D Models and 2D Images for Recognition and ClassificationabstractSimilarity measurements between 3D objects and 2D images are useful for the tasks of object recognition and classification. The authors distinguish between two types of similarity metrics: metrics computed in image-space (image metrics) and metrics computed in transformation-space (transformation metrics). Existing methods typically use image metrics; namely, metrics that measure the difference in the image between the observed image and the nearest view of the object. Example for such a measure is the Euclidean distance between feature points in the image and their corresponding points in the nearest view. (This measure can be computed by solving the exterior orientation calibration problem.) In this paper the authors introduce a different type of metrics: transformation metrics. These metrics penalize for the deformations applied to the object to produce the observed image. In particular, the authors define a transformation metric that optimally penalizes for "affine deformations" under weak-perspective. A closed-form solution, together with the nearest view according to this metric, are derived. The metric is shown to be equivalent to the Euclidean image metric, in the sense that they bound each other from both above and below. It therefore provides an easy-to-use closed-form approximation for the commonly-used least-squares distance between models and images. The authors demonstrate an image understanding application, where the true dimensions of a photographed battery charger are estimated by minimizing the transformation metric. Ronen Basri, Daphna Weinshall |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | Recognition Using Region CorrespondencesabstractA central problem in object recognition is to determine the transformation that relates the model to the image, given some partial correspondence between the two. This is useful in determining whether an object is present in an image, and if so, in determining where the object is. We present a novel method of solving this problem that uses region information. In our approach, the model is divided into volumes and the image is divided into regions. Given a match between subsets of volumes and regions (without any explicit correspondence between different pieces of the regions), the alignment transformation is computed. The method applies to planar objects under similarity, affine and projective transformations and to projections of 3D objects undergoing affine and projective transformations.> Ronen Basri, David Jacobs 0001 |
ICCV | 1 |
| 1995 | Localization and Homing Using Combinations of Model Views
Ronen Basri, Ehud Rivlin |
Artif. Intell. | 1 |
| 1994 | Navigation based on a network of 2D imagesabstractThis paper describes the integration of 2D stimulus-driven robot localization and positioning with a token-based correspondence method in a practical robot navigation system. The approach allows for modular acquisition and update of world knowledge for navigation, and robustness of navigation to low-level errors. No special marking of the world is necessary, so the robot may operate in quite general environments. Tests in a real industrial environment confirm the potential of the method. David Wilkes, Sven J. Dickinson, Ehud Rivlin, Ronen Basri |
ICPR (1) | 4 |
| 1993 | Recognition by prototypesabstractA scheme for recognizing 3D objects from single 2D images is introduced. The scheme proceeds in two stages. In the categorization stage, the image is matched against prototype objects, and in the identification stage, the observed object is matched against the individual models of its class, where classes are expected to contain objects with relatively similar shapes. The advantage of categorizing the object before it is identified is twofold. First, the image is compared to a smaller number of models, since only models that belong to the object's class need to be considered. Second, the cost of comparing the image to each model in a class is very low, because correspondence is computed once for the whole class. The correspondence and object pose computed in the categorization stage to align the prototype with the image are reused in the identification stage to align the individual models with the image. As a result, identification is reduced to a series of simple template comparisons.> Ronen Basri |
CVPR | 1 |
| 1993 | Distance metric between 3D models and 2D images for recognition and classificationabstractA transformation metric to measure the similarity between 3-D models and 2-D images is proposed. The transformation metric measures the amount of affine deformation applied to the object to produce the given image. A simple, closed-form solution for this metric is presented. This solution is optimal in transformation space, and it is used to bound the image metric from both above and below. The transformation metric can be used in several different ways in recognition and classification tasks.> Daphna Weinshall, Ronen Basri |
CVPR | 2 |
| 1993 | Localization using combinations of model viewsabstractA method for localization, the act of recognizing the environment, is presented. The method is based on representing the scene as a set of 2-D views and predicting the appearances of novel views by linear combinations of the model views. The method accurately approximates the appearance of scenes under weak perspective projection. Analysis of this projection as well as experimental results demonstrate that in many cases this approximation is sufficient to accurately describe the scene. When weak perspective approximation is invalid, either a larger number of models can be acquired or an iterative solution to account for the perspective distortions can be used. The method has several advantages over other approaches. It uses relatively rich representations; the representations are 2-D rather than 3-D; and localization can be done from only a single 2-D view.> Ronen Basri, Ehud Rivlin |
ICCV | 1 |
| 1993 | Homing Using Combinations of Model Views
Ronen Basri, Ehud Rivlin |
IJCAI | 1 |
| 1992 | The alignment of objects with smooth surfaces: error analysis of the curvature methodabstractThe recognition of objects with smooth bounding surfaces from their contour images is addressed. In particular, the curvature method is applied to ellipsoidal objects and the error for different rotations of the objects is computed analytically. It is seen that the error depends on the exact shape of the ellipsoid (namely, the relative lengths of its axes), and it increases as the ellipsoid becomes elongated in the Z-direction. It is shown that the errors are usually small, and that, in general, a small number of models is required to predict the appearance of an ellipsoid from all possible views.> Ronen Basri |
CVPR | 1 |
| 1991 | Linear Operator for Object Recognition
Ronen Basri, Shimon Ullman |
NIPS | 1 |
| 1991 | Recognition by Linear Combinations of ModelsabstractAn approach to visual object recognition in which a 3D object is represented by the linear combination of 2D images of the object is proposed. It is shown that for objects with sharp edges as well as with smooth bounding contours, the set of possible images of a given object is embedded in a linear space spanned by a small number of views. For objects with sharp edges, the linear combination representation is exact. For objects with smooth boundaries, it is an approximation that often holds over a wide range of viewing angles. Rigid transformations (with or without scaling) can be distinguished from more general linear transformations of the object by testing certain constraints placed on the coefficients of the linear combinations. Three alternative methods of determining the transformation that matches a model to a given image are proposed.> Shimon Ullman, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1988 | The Alignment Of Objects With Smooth SurfacesabstractAbstract : This paper examines the recognition of rigid objects bounded by smooth surfaces, using an alignment approach. The projected image of such an object changes during rotation in a manner that is generally difficult to predict. An approach to this problem is suggested, using the 3-D surface curvature at the points along the silhouette. The curvature information requires a single number for each point along the object's silhouette, the magnitude of the curvature vector at the point. We have implemented this method, and tested it on images of complex 3-D objects. Models of the viewed objects were acquired using three images of each object. The implemented scheme was found to give accurate predictions of the objects' appearance for large transformations. Using this method, a small number of (viewer-centered) models can be used to predict the new appearance of an object from any given viewpoint. (JHD) Ronen Basri, Shimon Ullman |
ICCV | 1 |