Tomás Pajdla

dblp:p/TomasPajdla · DBLP profile ↗
← Back
138ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0001-6325-0072ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 123 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 97 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Algebraic Geometry Approach to Viewing Graph Solvability
abstract
The concept of viewing graph solvability has gained significant interest in the context of structure-from-motion. A viewing graph is a mathematical structure where nodes are associated with cameras and edges represent the epipolar geometry connecting overlapping views. Solvability studies under which conditions the cameras are uniquely determined by the graph. In this paper we propose a novel framework for analyzing solvability problems based on algebraic geometry, demonstrating its potential in understanding structure-from-motion graphs and proving a conjecture that was previously proposed.
Federica Arrigoni, Kathlén Kohn, Andrea Fusiello, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Order-One Rolling Shutter Cameras
abstract
Rolling shutter (RS) cameras dominate consumer and smartphone markets. Several methods for computing the absolute pose of RS cameras have appeared in the last 20 years, but the relative pose problem has not been fully solved yet. We provide a unified theory for the important class of order-one rolling shutter (RS1) cameras. These cameras generalize the perspective projection to RS cameras, projecting a generic space point to exactly one image point via a rational map. We introduce a new back-projection RS camera model, characterize RS1cameras, construct explicit parameterizations of such cameras, and determine the image of a space line. We classify all minimal problems for solving the relative camera pose problem with linear RS1cameras and discover new practical cases. Finally, we show how the theory can be used to explain RS models previously used for absolute pose computation.
Marvin Anas Hahn, Kathlén Kohn, Orlando Marigliano, Tomás Pajdla
CVPR4
2025 Revisiting Viewing Graph Solvability: An Effective Approach Based on Cycle Consistency
abstract
In the structure from motion, the viewing graph is a graph where the vertices correspond to cameras (or images) and the edges represent the fundamental matrices. We provide a new formulation and an algorithm for determining whether a viewing graph is solvable, i.e., uniquely determines a set of projective cameras. The known theoretical conditions either do not fully characterize the solvability of all viewing graphs, or are extremely difficult to compute because they involve solving a system of polynomial equations with a large number of unknowns. The main result of this paper is a method to reduce the number of unknowns by exploiting cycle consistency. We advance the understanding of solvability by (i) finishing the classification of all minimal graphs up to 9 nodes, (ii) extending the practical verification of solvability to minimal graphs with up to 90 nodes, (iii) finally answering an open research question by showing that finite solvability is not equivalent to solvability, and (iv) formally drawing the connection with the calibrated case (i.e., parallel rigidity). Finally, we present an experiment on real data that shows that unsolvable graphs may appear in practice.
Federica Arrigoni, Andrea Fusiello, Romeo Rizzi, Elisa Ricci 0001, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Learning to Solve Hard Minimal Problems
Petr Hruby, Timothy Duff, Anton Leykin, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Minimal Perspective Autocalibration
abstract
We introduce a new family of minimal problems for reconstruction from multiple views. Our primary focus is a novel approach to autocalibration, a long-standing problem in computer vision. Traditional approaches to this problem, such as those based on Kruppa's equations or the modulus constraint, rely explicitly on the knowledge of multiple fundamental matrices or a projective reconstruction. In contrast, we consider a novel formulation involving constraints on image points, the unknown depths of 3D points, and a partially specified calibration matrix$K$. For 2 and 3 views, we present a comprehensive taxonomy of minimal autocalibration problems obtained by relaxing some of these constraints. These problems are organized into classes according to the number of views and any assumed prior knowledge of$K$. Within each class, we determine problems with the fewest—or a relatively small number of—solutions. From this zoo of problems, we devise three practical solvers. Experiments with synthetic and real data and interfacing our solvers with COLMAP demonstrate that we achieve superior accuracy compared to state-of-the-art calibration methods. The code is available at github.com/andreadalcin/MinimalPerspectiveAutocalibration.
Andrea Porfiri Dal Cin, Timothy Duff, Luca Magri 0002, Tomás Pajdla
CVPR4
2024 A Direct Approach to Viewing Graph Solvability
Federica Arrigoni, Andrea Fusiello, Tomás Pajdla
ECCV (1)3
2024 PL1P: Point-Line Minimal Problems under Partial Visibility in Three Views
Timothy Duff, Kathlén Kohn, Anton Leykin, Tomás Pajdla
Int. J. Comput. Vis.4
2024 Guest Editorial: Special Issue on Traditional Computer Vision in the Age of Deep Learning
Matteo Poggi, Federica Arrigoni, Andrea Fusiello, Stefano Mattoccia, Adrien Bartoli, Torsten Sattler, Tomás Pajdla
Int. J. Comput. Vis.7
2024 PLMP - Point-Line Minimal Problems in Complete Multi-View Visibility
abstract
We present a complete classification of all minimal problems for generic arrangements of points and lines completely observed by calibrated perspective cameras. We show that there are only 30 minimal problems in total, no problems exist for more than 6 cameras, for more than 5 points, and for more than 6 lines. We present a sequence of tests for detecting minimality starting with counting degrees of freedom and ending with full symbolic and numeric verification of representative examples. For all minimal problems discovered, we present their algebraic degrees, i.e.the number of solutions, which measure their intrinsic difficulty. It shows how exactly the difficulty of problems grows with the number of views. Importantly, several new minimal problems have small degrees that might be practical in image matching and 3D reconstruction.
Timothy Duff, Kathlén Kohn, Anton Leykin, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Four-view Geometry with Unknown Radial Distortion
abstract
We present novel solutions to previously unsolved prob-lems of relative pose estimation from images whose calibration parameters, namely focal lengths and radial distortion, are unknown. Our approach enables metric reconstruction without modeling these parameters. The minimal case for reconstruction requires 13 points in 4 views for both the calibrated and uncalibrated cameras. We describe and implement the first solution to these minimal problems. In the calibrated case, this may be modeled as a polynomial sys-tem of equations with 3584 solutions. Despite the apparent intractability, the problem decomposes spectacularly. Each solution falls into a Euclidean symmetry class of size 16, and we can estimate 224 class representatives by solving a sequence of three subproblems with 28, 2, and 4 solutions. We highlight the relationship between internal constraints on the radial quadrifocal tensor and the relations among the principal minors of a$4\times 4$matrix. We also address the case of 4 upright cameras, where 7 points are minimal. Finally, we evaluate our approach on simulated and real data and benchmark against previous calibration-free solutions, and show that our method provides an efficient startup for an SfM pipeline with radial cameras.
Petr Hruby, Viktor Korotynskiy, Timothy Duff, Luke Oeding, Marc Pollefeys, Tomás Pajdla, Viktor Larsson
CVPR6
2023 Viewing Graph Solvability in Practice
abstract
We present an advance in understanding the projective Structure-from-Motion, focusing in particular on the viewing graph: such a graph has cameras as nodes and fundamental matrices as edges. We propose a practical method for testing finite solvability, i.e., whether a viewing graph induces a finite number of camera configurations. Our formulation uses a significantly smaller number of equations (up to 400×) with respect to previous work. As a result, this is the only method in the literature that can be applied to large viewing graphs coming from real datasets, comprising up to 300K edges. In addition, we develop the first algorithm for identifying maximal finite-solvable components.
Federica Arrigoni, Tomás Pajdla, Andrea Fusiello
ICCV2
2023 Using monodromy to recover symmetries of polynomial systems
abstract
Galois/monodromy groups attached to parametric systems of polynomial equations provide a method for detecting the existence of symmetries in solution sets. Beyond the question of existence, one would like to compute formulas for these symmetries, towards the eventual goal of solving the systems more efficiently. We describe and implement one possible approach to this task using numerical homotopy continuation and multivariate rational function interpolation. We illustrate our methods on several examples, including two cases with nonlinear symmetries which appear in applications from computer vision and robotics.
Timothy Duff, Viktor Korotynskiy, Tomás Pajdla, Margaret H. Regan
ISSAC3
2023 Guest Editorial: Special Issue on Advances in Computer Vision and Applications (ACCV 2020)
Hiroshi Ishikawa 0002, Tomás Pajdla, Jianbo Shi
Int. J. Comput. Vis.3
2023 An Efficient Model for a Camera Behind a Parallel Refractive Slab
Aless Lasaruk, Tomás Pajdla
Int. J. Comput. Vis.2
2023 Trifocal Relative Pose From Lines at Points
abstract
We present a method for solving two minimal problems for relative camera pose estimation from three views, which are based on three view correspondences of (i) three points and one line and the novel case of (ii) three points and two lines through two of the points. These problems are too difficult to be efficiently solved by the state of the art Gröbner basis methods. Our method is based on a new efficient homotopy continuation (HC) solver framework MINUS, which dramatically speeds up previous HC solving by specializing hc methods to generic cases of our problems. We characterize their number of solutions and show with simulated experiments that our solvers are numerically robust and stable under image noise, a key contribution given the borderline intractable degree of nonlinearity of trinocular constraints. We show in real experiments that (i) sift feature location and orientation provide good enough point-and-line correspondences for three-view reconstruction and (ii) that we can solve difficult cases with too few or too noisy tentative matches, where the state of the art structure from motion initialization fails.
Ricardo Fabbri, Timothy Duff, Hongyi Fan, Margaret H. Regan, David da Costa de Pinho, Elias P. Tsigaridas, Charles W. Wampler, Jonathan D. Hauenstein, Peter J. Giblin, Benjamin B. Kimia, Anton Leykin, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.12
2022 Learning to Solve Hard Minimal Problems
abstract
We present an approach to solving hard geometric optimization problems in the RANSAC framework. The hard minimal problems arise from relaxing the original geometric optimization problem into a minimal problem with many spurious solutions. Our approach avoids computing large numbers of spurious solutions. We design a learning strategy for selecting a starting problem-solution pair that can be numerically continued to the problem and the solution of interest. We demonstrate our approach by developing a RANSAC solver for the problem of computing the relative pose of three calibrated cameras, via a minimal relaxation using four points in each view. On average, we can solve a single problem in under 70$\mu s.$μs. We also benchmark and study our engineering choices on the very familiar problem of computing the relative pose of two calibrated cameras, via the minimal case of five points in two views.
Petr Hruby, Timothy Duff, Anton Leykin, Tomás Pajdla
CVPR4
2022 Optimizing Elimination Templates by Greedy Parameter Search
abstract
We propose a new method for constructing elimination templates for efficient polynomial system solving of minimal problems in structure from motion, image matching, and camera tracking. We first construct a particular affine parameterization of the elimination templates for systems with a finite number of distinct solutions. Then, we use a heuristic greedy optimization strategy over the space of parameters to get a template with a small size. We test our method on 34 minimal problems in computer vision. For all of them, we found the templates either of the same or smaller size compared to the state-of-the-art. For some difficult examples, our templates are, e.g., 2.1, 2.5, 3.8, 6.6 times smaller. For the problem of refractive absolute pose estimation with unknown focal length, we have found a template that is 20 times smaller. Our experiments on synthetic data also show that the new solvers are fast and numerically accurate. We also present a fast and numerically accurate solver for the problem of relative pose estimation with unknown common focal length and radial distortion.
Evgeniy Martyushev, Jana Vráblíková, Tomás Pajdla
CVPR3
2022 Objects Can Move: 3D Change Detection by Geometric Transformation Consistency
Aikaterini Adam, Torsten Sattler, Konstantinos Karantzalos, Tomás Pajdla
ECCV (33)4
2022 Multi-frame Motion Segmentation by Combining Two-Frame Results
abstract
Abstract In this paper we consider the motion segmentation problem on sparse and unstructured datasets involving rigid motions, motivated by multibody structure from motion. In particular, we assume only two-frame correspondences as input without prior knowledge about trajectories. Inspired by the success of synchronization methods, we address this problem by introducing a two-stage approach: first, motion segmentation is addressed on image pairs independently; then, two-frame results are combined in a robust way to compute the final multi-frame segmentation. Our synthetic and real experiments demonstrate that the proposed approach is very effective in reducing the errors among two-frame results and it can cope with a large amount of mismatches. Moreover, our method can be profitably used to build a multibody structure from motion pipeline.
Federica Arrigoni, Elisa Ricci 0001, Tomás Pajdla
Int. J. Comput. Vis.3
2022 NCNet: Neighbourhood Consensus Networks for Estimating Image Correspondences
abstract
We address the problem of finding reliable dense correspondences between a pair of images. This is a challenging task due to strong appearance differences between the corresponding scene elements and ambiguities generated by repetitive patterns. The contributions of this work are threefold. First, inspired by the classic idea of disambiguating feature matches using semi-local constraints, we develop an end-to-end trainable convolutional neural network architecture that identifies sets of spatially consistent matches by analyzing neighbourhood consensus patterns in the 4D space of all possible correspondences between a pair of images without the need for a global geometric model. Second, we demonstrate that the model can be trained effectively from weak supervision in the form of matching and non-matching image pairs without the need for costly manual annotation of point to point correspondences. Third, we show the proposed neighbourhood consensus network can be applied to a range of matching tasks including both category- and instance-level matching, obtaining the state-of-the-art results on the PF, TSS, InLoc, and HPatches benchmarks.
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Akihiko Torii, Tomás Pajdla, Josef Sivic
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Long-Term Visual Localization Revisited
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing conditions, including day-night changes, as well as weather and seasonal variations, while providing highly accurate six degree-of-freedom (6DOF) camera pose estimates. In this paper, we extend three publicly available datasets containing images captured under a wide variety of viewing conditions, but lacking camera pose information, with ground truth pose information, making evaluation of the impact of various factors on 6DOF camera pose estimation accuracy possible. We also discuss the performance of state-of-the-art localization approaches on these datasets. Additionally, we release around half of the poses for all conditions, and keep the remaining half private as a test set, in the hopes that this will stimulate research on long-term visual localization, learned local image features, and related research areas. Our datasets are available at visuallocalization.net, where we are also hosting a benchmarking server for automatic evaluation of results on the test set. The presented state-of-the-art results are to a large degree based on submissions to our server.
Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Fredrik Kahl, Torsten Sattler
IEEE Trans. Pattern Anal. Mach. Intell.10
2021 Viewing Graph Solvability via Cycle Consistency
abstract
In structure-from-motion the viewing graph is a graph where vertices correspond to cameras and edges represent fundamental matrices. We provide a new formulation and an algorithm for establishing whether a viewing graph is solvable, i.e. it uniquely determines a set of projective cameras. Known theoretical conditions either do not fully characterize the solvability of all viewing graphs, or are exceedingly hard to compute for they involve solving a system of polynomial equations with a large number of unknowns. The main result of this paper is a method for reducing the number of unknowns by exploiting the cycle consistency. We advance the understanding of the solvability by (i) finishing the classification of all previously undecided minimal graphs up to 9 nodes, (ii) extending the practical solvability testing up to minimal graphs with up to 90 nodes, and (iii) definitely answering an open research question by showing that the finite solvability is not equivalent to the solvability. Finally, we present an experiment on real data showing that unsolvable graphs are appearing in practical situations.
Federica Arrigoni, Andrea Fusiello, Elisa Ricci 0001, Tomás Pajdla
ICCV4
2021 InLoc: Indoor Visual Localization with Dense Matching and View Synthesis
abstract
We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor spaces. The method proceeds along three steps: (i) efficient retrieval of candidate poses that scales to large-scale environments, (ii) pose estimation using dense matching rather than sparse local features to deal with weakly textured indoor scenes, and (iii) pose verification by virtual view synthesis that is robust to significant changes in viewpoint, scene layout, and occlusion. Second, we release a new dataset with reference 6DoF poses for large-scale indoor localization. Query photographs are captured by mobile phones at a different time than the reference 3D map, thus presenting a realistic indoor localization scenario. Third, we demonstrate that our method significantly outperforms current state-of-the-art indoor localization approaches on this new challenging data. Code and data are publicly available.
Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Akihiko Torii
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?
abstract
Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6 degree-of-freedom (DOF) pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast, 2D image retrieval-based methods only require a database of geo-tagged images, which is trivial to construct and to maintain. They are often considered inaccurate since they only approximate the positions of the cameras. Yet, the exact camera pose can theoretically be recovered when enough relevant database images are retrieved. In this paper, we demonstrate experimentally that large-scale 3D models are not strictly necessary for accurate visual localization. We create reference poses for a large and challenging urban dataset. Using these poses, we show that combining image-based methods with local reconstructions results in a higher pose accuracy compared to state-of-the-art structure-based methods, albeight at higher run-time costs. We show that some of these run-time costs can be alleviated by exploiting known database image poses. Our results suggest that we might want to reconsider the need for large-scale 3D models in favor of more local models, but also that further research is necessary to accelerate the local reconstruction process.
Akihiko Torii, Hajime Taira, Josef Sivic, Marc Pollefeys, Masatoshi Okutomi, Tomás Pajdla, Torsten Sattler
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 From Two Rolling Shutters to One Global Shutter
abstract
Most consumer cameras are equipped with electronic rolling shutter, leading to image distortions when the camera moves during image capture. We explore a surprisingly simple camera configuration that makes it possible to undo the rolling shutter distortion: two cameras mounted to have different rolling shutter directions. Such a setup is easy and cheap to build and it possesses the geometric constraints needed to correct rolling shutter distortion using only a sparse set of point correspondences between the two images. We derive equations that describe the underlying geometry for general and special motions and present an efficient method for finding their solutions. Our synthetic and real experiments demonstrate that our approach is able to remove large rolling shutter distortions of all types without relying on any specific scene structure.
Cenek Albl, Zuzana Kukelova, Viktor Larsson, Michal Polic, Tomás Pajdla, Konrad Schindler
CVPR5
2020 TRPLP - Trifocal Relative Pose From Lines at Points
abstract
We present a method for solving two minimal problems for relative camera pose estimation from three views, which are based on three view correspondences of (i) three points and one line and (ii) three points and two lines through two of the points. These problems are too difficult to be efficiently solved by the state of the art Grobner basis methods. Our method is based on a new efficient homotopy continuation (HC) solver, which dramatically speeds up previous HC solving by specializing HC methods to generic cases of our problems. We show in simulated experiments that our solvers are numerically robust and stable under image noise. We show in real experiment that (i) SIFT features provide good enough point-and-line correspondences for three-view reconstruction and (ii) that we can solve difficult cases with too few or too noisy tentative matches where the state of the art structure from motion initialization fails.
Ricardo Fabbri, Timothy Duff, Hongyi Fan, Margaret H. Regan, David da Costa de Pinho, Elias P. Tsigaridas, Charles W. Wampler, Jonathan D. Hauenstein, Peter J. Giblin, Benjamin B. Kimia, Anton Leykin, Tomás Pajdla
CVPR12
2020 Uncertainty Based Camera Model Selection
abstract
The quality and speed of Structure from Motion (SfM) methods depend significantly on the camera model chosen for the reconstruction. In most of the SfM pipelines, the camera model is manually chosen by the user. In this paper, we present a new automatic method for camera model selection in large scale SfM that is based on efficient uncertainty evaluation. We first perform an extensive comparison of classical model selection based on known Information Criteria and show that they do not provide sufficiently accurate results when applied to camera model selection. Then we propose a new Accuracy-based Criterion, which evaluates an efficient approximation of the uncertainty of the estimated parameters in tested models. Using the new criterion, we design a camera model selection method and fine-tune it by machine learning. Our simulated and real experiments demonstrate a significant increase in reconstruction quality as well as a considerable speedup of the SfM process.
Michal Polic, Stanislav Steidl, Cenek Albl, Zuzana Kukelova, Tomás Pajdla
CVPR5
2020 On the Usage of the Trifocal Tensor in Motion Segmentation
Federica Arrigoni, Luca Magri 0002, Tomás Pajdla
ECCV (20)3
2020 Making Affine Correspondences Work in Camera Geometry Computation
Daniel Barath, Michal Polic, Wolfgang Förstner, Torsten Sattler, Tomás Pajdla, Zuzana Kukelova
ECCV (11)5
2020 PL1P - Point-Line Minimal Problems Under Partial Visibility in Three Views
Timothy Duff, Kathlén Kohn, Anton Leykin, Tomás Pajdla
ECCV (26)4
2020 Minimal Rolling Shutter Absolute Pose with Unknown Focal Length and Radial Distortion
Zuzana Kukelova, Cenek Albl, Akihiro Sugimoto, Konrad Schindler, Tomás Pajdla
ECCV (5)5
2020 Motion Segmentation with Pairwise Matches and Unknown Number of Motions
abstract
In this paper we address motion segmentation, that is the problem of clustering points in multiple images according to a number of moving objects. Two-frame correspondences are assumed as input without prior knowledge about trajectories. Our method is based on principles from “multi-model fitting” and “permutation synchronization”, and - differently from previous techniques working under the same assumptions - it can handle an unknown number of motions. The proposed approach is validated on standard datasets, showing that it can correctly estimate the number of motions while maintaining comparable or better accuracy than the state of the art.
Federica Arrigoni, Luca Magri 0002, Tomás Pajdla
ICPR3
2020 Rolling Shutter Camera Absolute Pose
abstract
We present minimal, non-iterative solutions to the absolute pose problem for images from rolling shutter cameras. The absolute pose problem is a key problem in computer vision and rolling shutter is present in a vast majority of today's digital cameras. We discuss several camera motion models and propose two feasible rolling shutter camera models for a polynomial solver. In previous work a linearized camera model was used that required an initial estimate of the camera orientation. We show how to simplify the system of equations and make this solver faster. Furthermore, we present a first solution of the non-linearized camera orientation model using the Cayley parameterization. The new solver does not require any initial camera orientation estimate and therefore serves as a standalone solution to the rolling shutter camera pose problem from six 2D-to-3D correspondences. We show that our algorithms outperform P3P followed by a non-linear refinement using a rolling shutter model.
Cenek Albl, Zuzana Kukelova, Viktor Larsson, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 D2-Net: A Trainable CNN for Joint Description and Detection of Local Features
abstract
In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the detection to a later stage, the obtained keypoints are more stable than their traditional counterparts based on early detection of low-level structures. We show that this model can be trained using pixel correspondences extracted from readily available large-scale SfM reconstructions, without any further annotations. The proposed method obtains state-of-the-art performance on both the difficult Aachen Day-Night localization dataset and the InLoc indoor localization benchmark, as well as competitive performance on other benchmarks for image matching and 3D reconstruction.
Mihai Dusmanu, Ignacio Rocco, Tomás Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, Torsten Sattler
CVPR3
2019 Robust Motion Segmentation From Pairwise Matches
abstract
In this paper we consider the problem of motion segmentation, where only pairwise correspondences are assumed as input without prior knowledge about tracks. The problem is formulated as a two-step process. First, motion segmentation is performed on image pairs independently. Secondly, we combine independent pairwise segmentation results in a robust way into the final globally consistent segmentation. Our approach is inspired by the success of averaging methods. We demonstrate in simulated as well as in real experiments that our method is very effective in reducing the errors in the pairwise motion segmentation and can cope with large number of mismatches.
Federica Arrigoni, Tomás Pajdla
ICCV2
2019 PLMP - Point-Line Minimal Problems in Complete Multi-View Visibility
abstract
We present a complete classification of all minimal problems for generic arrangements of points and lines completely observed by calibrated perspective cameras. We show that there are only 30 minimal problems in total, no problems exist for more than 6 cameras, for more than 5 points, and for more than 6 lines. We present a sequence of tests for detecting minimality starting with counting degrees of freedom and ending with full symbolic and numeric verification of representative examples. For all minimal problems discovered, we present their algebraic degrees, i.e. the number of solutions, which measure their intrinsic difficulty. It shows how exactly the difficulty of problems grows with the number of views. Importantly, several new mini- mal problems have small degrees that might be practical in image matching and 3D reconstruction.
Timothy Duff, Kathlén Kohn, Anton Leykin, Tomás Pajdla
ICCV4
2019 Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization
abstract
Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common strategy is, first, to generate multiple estimates for the camera pose from which a given query image was taken. The pose with the largest geometric consistency with the query image, e.g., in the form of an inlier count, is then selected in a second stage. While a significant amount of research has concentrated on the first stage, there has been considerably less work on the second stage. In this paper, we thus focus on pose verification. We show that combining different modalities, namely appearance, geometry, and semantics, considerably boosts pose verification and consequently pose accuracy. We develop multiple hand-crafted as well as a trainable approach to join into the geometric-semantic verification and show significant improvements over state-of-the-art on a very challenging indoor dataset.
Hajime Taira, Ignacio Rocco, Jirí Sedlár, Masatoshi Okutomi, Josef Sivic, Tomás Pajdla, Torsten Sattler, Akihiko Torii
ICCV6
2018 Linear Solution to the Minimal Absolute Pose Rolling Shutter Problem
Zuzana Kukelova, Cenek Albl, Akihiro Sugimoto, Tomás Pajdla
ACCV (3)4
2018 Beyond Grobner Bases: Basis Selection for Minimal Solvers
abstract
Many computer vision applications require robust estimation of the underlying geometry, in terms of camera motion and 3D structure of the scene. These robust methods often rely on running minimal solvers in a RANSAC framework. In this paper we show how we can make polynomial solvers based on the action matrix method faster, by careful selection of the monomial bases. These monomial bases have traditionally been based on a Grobner basis for the polynomial ideal. Here we describe how we can enumerate all such bases in an efficient way. We also show that going beyond Grobner bases leads to more efficient solvers in many cases. We present a novel basis sampling scheme that we evaluate on a number of problems.
Viktor Larsson, Magnus Oskarsson, Kalle Åström, Alge Wallis, Zuzana Kukelova, Tomás Pajdla
CVPR6
2018 Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and seasonal variations, while providing highly accurate 6 degree-of-freedom (6DOF) camera pose estimates. In this paper, we introduce the first benchmark datasets specifically designed for analyzing the impact of such factors on visual localization. Using carefully created ground truth poses for query images taken under a wide variety of conditions, we evaluate the impact of various factors on 6DOF camera pose estimation accuracy through extensive experiments with state-of-the-art localization approaches. Based on our results, we draw conclusions about the difficulty of different conditions, showing that long-term localization is far from solved, and propose promising avenues for future work, including sequence-based localization approaches and the need for better local features. Our benchmark is available at visuallocalization.net.
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, Tomás Pajdla
CVPR12
2018 InLoc: Indoor Visual Localization With Dense Matching and View Synthesis
abstract
We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor environments. The method proceeds along three steps: (i) efficient retrieval of candidate poses that ensures scalability to large-scale environments, (ii) pose estimation using dense matching rather than local features to deal with texture less indoor scenes, and (iii) pose verification by virtual view synthesis to cope with significant changes in viewpoint, scene layout, and occluders. Second, we collect a new dataset with reference 6DoF poses for large-scale indoor localization. Query photographs are captured by mobile phones at a different time than the reference 3D map, thus presenting a realistic indoor localization scenario. Third, we demonstrate that our method significantly outperforms current state-of-the-art indoor localization approaches on this new challenging data.
Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Akihiko Torii
CVPR7
2018 Fast and Accurate Camera Covariance Computation for Large 3D Reconstruction
Michal Polic, Wolfgang Förstner, Tomás Pajdla
ECCV (2)3
2018 Neighbourhood Consensus Networks
abstract
We address the problem of finding reliable dense correspondences between a pair of images. This is a challenging task due to strong appearance differences between the corresponding scene elements and ambiguities generated by repetitive patterns. The contributions of this work are threefold. First, inspired by the classic idea of disambiguating feature matches using semi-local constraints, we develop an end-to-end trainable convolutional neural network architecture that identifies sets of spatially consistent matches by analyzing neighbourhood consensus patterns in the 4D space of all possible correspondences between a pair of images without the need for a global geometric model. Second, we demonstrate that the model can be trained effectively from weak supervision in the form of matching and non-matching image pairs without the need for costly manual annotation of point to point correspondences. Third, we show the proposed neighbourhood consensus network can be applied to a range of matching tasks including both category- and instance-level matching, obtaining the state-of-the-art results on the PF Pascal dataset and the InLoc indoor visual localization benchmark.
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Akihiko Torii, Tomás Pajdla, Josef Sivic
NeurIPS5
2018 NetVLAD: CNN Architecture for Weakly Supervised Place Recognition
Relja Arandjelovic, Petr Gronát, Akihiko Torii, Tomás Pajdla, Josef Sivic
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 24/7 Place Recognition by View Synthesis
abstract
We address the problem of large-scale visual place recognition for situations where the scene undergoes a major change in appearance, for example, due to illumination (day/night), change of seasons, aging, or structural modifications over time such as buildings being built or destroyed. Such situations represent a major challenge for current large-scale place recognition methods. This work has the following three principal contributions. First, we demonstrate that matching across large changes in the scene appearance becomes much easier when both the query image and the database image depict the scene from approximately the same viewpoint. Second, based on this observation, we develop a new place recognition approach that combines (i) an efficient synthesis of novel views with (ii) a compact indexable image representation. Third, we introduce a new challenging dataset of 1,125 camera-phone query images of Tokyo that contain major changes in illumination (day, sunset, night) as well as structural changes in the scene. We demonstrate that the proposed approach significantly outperforms other large-scale place recognition techniques on this challenging data.
Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Camera Uncertainty Computation in Large 3D Reconstruction
abstract
In computer vision, large scale Structure from Motion pipelines do not often evaluate the quality of the reconstruction by error propagation from measurements to the estimated parameters. It is a numerically sensitive and computationally challenging process, which is not easy to implement in practice for large scenes. We present a new algorithm that increases the numerical precision of the uncertainty propagation. It works with millions of feature points, thousands of cameras and millions of 3D points on a single computer. We provide an experimental comparison of our approach, as well as of previous approaches, on accurate ground truth and demonstrate that our algorithm is practical.
Michal Polic, Tomás Pajdla
3DV2
2017 On the Two-View Geometry of Unsynchronized Cameras
abstract
We present new methods of simultaneously estimating camera geometry and time shift from video sequences from multiple unsynchronized cameras. Algorithms for simultaneous computation of a fundamental matrix or a homography with unknown time shift between images are developed. Our methods use minimal correspondence sets (eight for fundamental matrix and four and a half for homography) and therefore are suitable for robust estimation using RANSAC. Furthermore, we present an iterative algorithm that extends the applicability on sequences which are significantly unsynchronized, finding the correct time shift up to several seconds. We evaluated the methods on synthetic and wide range of real world datasets and the results show a broad applicability to the problem of camera synchronization.
Cenek Albl, Zuzana Kukelova, Andrew W. Fitzgibbon, Jan Heller, Matej Smíd, Tomás Pajdla
CVPR6
2017 A Clever Elimination Strategy for Efficient Minimal Solvers
abstract
We present a new insight into the systematic generation of minimal solvers in computer vision, which leads to smaller and faster solvers. Many minimal problem formulations are coupled sets of linear and polynomial equations where image measurements enter the linear equations only. We show that it is useful to solve such systems by first eliminating all the unknowns that do not appear in the linear equations and then extending solutions to the rest of unknowns. This can be generalized to fully non-linear systems by linearization via lifting. We demonstrate that this approach leads to more efficient solvers in three problems of partially calibrated relative camera pose computation with unknown focal length and/or radial distortion. Our approach also generates new interesting constraints on the fundamental matrices of partially calibrated cameras, which were not known before.
Zuzana Kukelova, Joe Kileel 0001, Bernd Sturmfels, Tomás Pajdla
CVPR4
2017 Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?
abstract
Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6DOF pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast, 2D image retrieval-based methods only require a database of geo-tagged images, which is trivial to construct and to maintain. They are often considered inaccurate since they only approximate the positions of the cameras. Yet, the exact camera pose can theoretically be recovered when enough relevant database images are retrieved. In this paper, we demonstrate experimentally that large-scale 3D models are not strictly necessary for accurate visual localization. We create reference poses for a large and challenging urban dataset. Using these poses, we show that combining image-based methods with local reconstructions results in a pose accuracy similar to the state-of-the-art structure-based methods. Our results suggest that we might want to reconsider the current approach for accurate large-scale localization.
Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, Tomás Pajdla
CVPR7
2017 Nautilus: recovering regional symmetry transformations for image editing
abstract
Natural images often exhibit symmetries that should be taken into account when editing them. In this paper we present Nautilus --- a method for automatically identifying symmetric regions in an image along with their corresponding symmetry transformations. We compute dense local similarity symmetry transformations using a novel variant of the Generalised PatchMatch algorithm that uses Metropolis-Hastings sampling. We combine and refine these local symmetries using an extended Lucas-Kanade algorithm to compute regional transformations and their spatial extents. Our approach produces dense estimates of complex symmetries that are combinations of translation, rotation, scale, and reflection under perspective distortion. This enables a number of automatic symmetry-aware image editing applications including inpainting, rectification, beautification, and segmentation, and we demonstrate state-of-the-art applications for each of them.
Michal Lukác, Daniel Sýkora, Kalyan Sunkavalli, Eli Shechtman, Ondrej Jamriska, Nathan Carr 0001, Tomás Pajdla
ACM Trans. Graph.7
2016 Rolling Shutter Absolute Pose Problem with Known Vertical Direction
abstract
We present a solution to the rolling shutter (RS) absolute camera pose problem with known vertical direction. Our new solver, R5Pup, is an extension of the general minimal solution R6P, which uses a double linearized RS camera model initialized by the standard perspective P3P. Here, thanks to using known vertical directions, we avoid double linearization and can get the camera absolute pose directly from the RS model without the initialization by a standard P3P. Moreover, we need only five 2D-to-3D matches while R6P needed six such matches. We demonstrate in simulated and real experiments that our new R5Pup is robust, fast and a very practical method for absolute camera pose computation for modern cameras on mobile devices. We compare our R5Pup to the state of the art RS and perspective methods and demonstrate that it outperforms them when vertical direction is known in the range of accuracy available on modern mobile devices. We also demonstrate that when using R5Pup solver in structure from motion (SfM) pipelines, it is better to transform already reconstructed scenes into the standard position, rather than using hard constraints on the verticality of up vectors.
Cenek Albl, Zuzana Kukelova, Tomás Pajdla
CVPR3
2016 NetVLAD: CNN Architecture for Weakly Supervised Place Recognition
abstract
We tackle the problem of large scale visual place recognition, where the task is to quickly and accurately recognize the location of a given query photograph. We present the following four principal contributions. First, we develop a convolutional neural network (CNN) architecture that is trainable in an end-to-end manner directly for the place recognition task. The main component of this architecture, NetVLAD, is a new generalized VLAD layer, inspired by the "Vector of Locally Aggregated Descriptors" image representation commonly used in image retrieval. The layer is readily pluggable into any CNN architecture and amenable to training via backpropagation. Second, we create a new weakly supervised ranking loss, which enables end-to-end learning of the architecture's parameters from images depicting the same places over time downloaded from Google Street View Time Machine. Third, we develop an efficient training procedure which can be applied on very large-scale weakly labelled tasks. Finally, we show that the proposed architecture and training procedure significantly outperform non-learnt image representations and off-the-shelf CNN descriptors on challenging place recognition and image retrieval benchmarks.
Relja Arandjelovic, Petr Gronát, Akihiko Torii, Tomás Pajdla, Josef Sivic
CVPR4
2016 Degeneracies in Rolling Shutter SfM
Cenek Albl, Akihiro Sugimoto, Tomás Pajdla
ECCV (5)3
2016 Editorial
Jingyi Ju, Bastian Goldlücke, Richard Szeliski, Tomás Pajdla
Comput. Vis. Image Underst.4
2016 Learning and Calibrating Per-Location Classifiers for Visual Place Recognition
Petr Gronát, Josef Sivic, Guillaume Obozinski, Tomás Pajdla
Int. J. Comput. Vis.4
2016 Globally Optimal Hand-Eye Calibration Using Branch-and-Bound
abstract
This paper introduces a novel solution to the hand-eye calibration problem. It uses camera measurements directly and, at the same time, requires neither prior knowledge of the external camera calibrations nor a known calibration target. Our algorithm uses branch-and-bound approach to minimize an objective function based on the epipolar constraint. Further, it employs Linear Programming to decide the bounding step of the algorithm.Our technique is able to recover both the unknown rotation and translation simultaneously and the solution is guaranteed to be globally optimal with respect to the L∞-norm.
Jan Heller, Michal Havlena, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 R6P - Rolling shutter absolute pose problem
abstract
We present a minimal, non-iterative solution to the absolute pose problem for images from rolling shutter cameras. Absolute pose problem is a key problem in computer vision and rolling shutter is present in a vast majority of today's digital cameras. We propose several rolling shutter camera models and verify their feasibility for a polynomial solver. A solution based on linearized camera model is chosen and verified in several experiments. We use a linear approximation to the camera orientation, which is meaningful only around the identity rotation. We show that the standard P3P algorithm is able to estimate camera orientation within 6 degrees for camera rotation velocity as high as 30deg/frame. Therefore we can use the standard P3P algorithm to estimate camera orientation and to bring the camera rotation matrix close to the identity. Using this solution, camera position, orientation, translational velocity and angular velocity can be computed using six 2D-to-3D correspondences, with orientation error under half a degree and relative position error under 2%. A significant improvement in terms of the number of inliers in RANSAC is demonstrated.
Cenek Albl, Zuzana Kukelova, Tomás Pajdla
CVPR3
2015 Radial distortion homography
abstract
The importance of precise homography estimation is often underestimated even though it plays a crucial role in various vision applications such as plane or planarity detection, scene degeneracy tests, camera motion classification, image stitching, and many more. Ignoring the radial distortion component in homography estimation-even for classical perspective cameras-may lead to significant errors or totally wrong estimates. In this paper, we fill the gap among the homography estimation methods by presenting two algorithms for estimating homography between two cameras with different radial distortions. Both algorithms can handle planar scenes as well as scenes where the relative motion between the cameras is a pure rotation. The first algorithm uses the minimal number of five image point correspondences and solves a nonlinear system of polynomial equations using Gröbner basis method. The second algorithm uses a non-minimal number of six image point correspondences and leads to a simple system of two quadratic equations in two unknowns and one system of six linear equations. The proposed algorithms are fast, stable, and can be efficiently used inside a RANSAC loop.
Zuzana Kukelova, Jan Heller, Martin Bujnak, Tomás Pajdla
CVPR4
2015 24/7 place recognition by view synthesis
abstract
We address the problem of large-scale visual place recognition for situations where the scene undergoes a major change in appearance, for example, due to illumination (day/night), change of seasons, aging, or structural modifications over time such as buildings built or destroyed. Such situations represent a major challenge for current large-scale place recognition methods. This work has the following three principal contributions. First, we demonstrate that matching across large changes in the scene appearance becomes much easier when both the query image and the database image depict the scene from approximately the same viewpoint. Second, based on this observation, we develop a new place recognition approach that combines (i) an efficient synthesis of novel views with (ii) a compact indexable image representation. Third, we introduce a new challenging dataset of 1,125 camera-phone query images of Tokyo that contain major changes in illumination (day, sunset, night) as well as structural changes in the scene. We demonstrate that the proposed approach significantly outperforms other large-scale place recognition techniques on this challenging data.
Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
CVPR5
2015 Efficient Solution to the Epipolar Geometry for Radially Distorted Cameras
abstract
The estimation of the epipolar geometry of two cameras from image matches is a fundamental problem of computer vision with many applications. While the closely related problem of estimating relative pose of two different uncalibrated cameras with radial distortion is of particular importance, none of the previously published methods is suitable for practical applications. These solutions are either numerically unstable, sensitive to noise, based on a large number of point correspondences, or simply too slow for real-time applications. In this paper, we present a new efficient solution to this problem that uses 10 image correspondences. By manipulating ten input polynomial equations, we derive a degree 10 polynomial equation in one variable. The solutions to this equation are efficiently found using the Sturm sequences method. In the experiments, we show that the proposed solution is stable, noise resistant, and fast, and as such efficiently usable in a practical Structure-from-Motion pipeline.
Zuzana Kukelova, Jan Heller, Martin Bujnak, Andrew W. Fitzgibbon, Tomás Pajdla
ICCV5
2015 Visual Place Recognition with Repetitive Structures
abstract
Repeated structures such as building facades, fences or road markings often represent a significant challenge for place recognition. Repeated structures are notoriously hard for establishing correspondences using multi-view geometry. They violate the feature independence assumed in the bag-of-visual-words representation which often leads to over-counting evidence and significant degradation of retrieval performance. In this work we show that repeated structures are not a nuisance but, when appropriately represented, they form an important distinguishing feature for many places. We describe a representation of repeated structures suitable for scalable retrieval and geometric verification. The retrieval is based on robust detection of repeated image structures and a suitable modification of weights in the bag-of-visual-word model. We also demonstrate that the explicit detection of repeated patterns is beneficial for robust visual word matching for geometric verification. Place recognition results are shown on datasets of street-level imagery from Pittsburgh and San Francisco demonstrating significant gains in recognition performance compared to the standard bag-of-visual-words baseline as well as the more recently proposed burstiness weighting and Fisher vector encoding.
Akihiko Torii, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 World-Base Calibration by Global Polynomial Optimization
abstract
This paper presents a novel solution to the world-base calibration problem. It is applicable in situations where a known calibration target is observed by a camera attached to the end effector of a robotic manipulator. The presented method works by minimizing geometrically meaningful error function based on image projections. Our formulation leads to a non-convex multivariate polynomial optimization problem of a constant size. However, we show how such a problem can be relaxed using linear matrix inequality (LMI) relaxations and effectively solved using Semi definite Programming. Although the technique of LMI relaxations guaranties only a lower bound on the global minimum of the original problem, it can provide a certificate of optimality in cases when the global minimum is reached. Indeed, we reached the global minimum for all calibration tasks in our experiments with both synthetic and real data. The experiments also show that the presented method is fast and noise resistant.
Jan Heller, Tomás Pajdla
3DV2
2014 Match Box: Indoor Image Matching via Box-Like Scene Estimation
abstract
Key point matching in images of indoor scenes traditionally employs features like SIFT, GIST and HOG. While those features work very well for two images related to each other by small camera transformations, we commonly observe a drop in performance for patches representing scene elements visualized from a very different perspective. Since increasing the space of considered local transformations for feature matching decreases their discriminative abilities, we propose a more global approach inspired by the recent success of monocular scene understanding. In particular we propose to reconstruct a box-like model of the scene from every single image and use it to rectify images before matching. We show that a monocular scene model reconstruction and rectification preceding standard feature matching significantly improves key point matching and dramatically improves reconstruction of difficult indoor scenes.
Filip Srajer, Alexander G. Schwing, Marc Pollefeys, Tomás Pajdla
3DV4
2014 Stable Radial Distortion Calibration by Polynomial Matrix Inequalities Programming
Jan Heller, Didier Henrion, Tomás Pajdla
ACCV (1)3
2014 Singly-Bordered Block-Diagonal Form for Minimal Problem Solvers
Zuzana Kukelova, Martin Bujnak, Jan Heller, Tomás Pajdla
ACCV (2)4
2014 Hand-eye and robot-world calibration by global polynomial optimization
abstract
The need to relate measurements made by a camera to a different known coordinate system arises in many engineering applications. Historically, it appeared for the first time in the connection with cameras mounted on robotic systems. This problem is commonly known as hand-eye calibration. In this paper, we present several formulations of hand-eye calibration that lead to multivariate polynomial optimization problems. We show that the method of convex linear matrix inequality (LMI) relaxations can be used to effectively solve these problems and to obtain globally optimal solutions. Further, we show that the same approach can be used for the simultaneous hand-eye and robot-world calibration. Finally, we validate the proposed solutions using both synthetic and real datasets.
Jan Heller, Didier Henrion, Tomás Pajdla
ICRA3
2013 Fast and Stable Algebraic Solution to L2 Three-View Triangulation
abstract
In this paper we provide a new fast and stable algebraic solution to the problem of L2triangulation from three views. We use Lagrange multipliers to formulate the search for the minima of the L2objective function subject to equality constraints. Interestingly, we show that by relaxing the triangulation such that we do not require a single point in 3D, we get, after a linear correction, a solver that is faster, more stable and practically as accurate as the state-of-the-art L2-optimal algebraic solvers [24, 7, 8, 9]. In our formulation, we obtain a system of eight polynomial equations in eight unknowns, which we solve using the Groebner basis method. We get less (31) solutions than was the number (47-66) of solutions obtained in [24, 7, 8, 9] and our solver is more robust than [8, 9] w.r.t. critical configurations. We evaluate the precision and speed of our solver on both synthetic and real datasets.
Zuzana Kukelova, Tomás Pajdla, Martin Bujnak
3DV2
2013 Learning and Calibrating Per-Location Classifiers for Visual Place Recognition
abstract
The aim of this work is to localize a query photograph by finding other images depicting the same place in a large geotagged image database. This is a challenging task due to changes in viewpoint, imaging conditions and the large size of the image database. The contribution of this work is two-fold. First, we cast the place recognition problem as a classification task and use the available geotags to train a classifier for each location in the database in a similar manner to per-exemplar SVMs in object recognition. Second, as only few positive training examples are available for each location, we propose a new approach to calibrate all the per-location SVM classifiers using only the negative examples. The calibration we propose relies on a significance measure essentially equivalent to the p-values classically used in statistical hypothesis testing. Experiments are performed on a database of 25,000 geotagged street view images of Pittsburgh and demonstrate improved place recognition accuracy of the proposed approach over the previous work.
Petr Gronát, Guillaume Obozinski, Josef Sivic, Tomás Pajdla
CVPR4
2013 Visual Place Recognition with Repetitive Structures
abstract
Repeated structures such as building facades, fences or road markings often represent a significant challenge for place recognition. Repeated structures are notoriously hard for establishing correspondences using multi-view geometry. Even more importantly, they violate the feature independence assumed in the bag-of-visual-words representation which often leads to over-counting evidence and significant degradation of retrieval performance. In this work we show that repeated structures are not a nuisance but, when appropriately represented, they form an important distinguishing feature for many places. We describe a representation of repeated structures suitable for scalable retrieval. It is based on robust detection of repeated image structures and a simple modification of weights in the bag-of-visual-word model. Place recognition results are shown on datasets of street-level imagery from Pittsburgh and San Francisco demonstrating significant gains in recognition performance compared to the standard bag-of-visual-words baseline and more recently proposed burstiness weighting.
Akihiko Torii, Josef Sivic, Tomás Pajdla, Masatoshi Okutomi
CVPR3
2013 Real-Time Solution to the Absolute Pose Problem with Unknown Radial Distortion and Focal Length
abstract
The problem of determining the absolute position and orientation of a camera from a set of 2D-to-3D point correspondences is one of the most important problems in computer vision with a broad range of applications. In this paper we present a new solution to the absolute pose problem for camera with unknown radial distortion and unknown focal length from five 2D-to-3D point correspondences. Our new solver is numerically more stable, more accurate, and significantly faster than the existing state-of-the-art minimal four point absolute pose solvers for this problem. Moreover, our solver results in less solutions and can handle larger radial distortions. The new solver is straightforward and uses only simple concepts from linear algebra. Therefore it is simpler than the state-of-the-art Groebner basis solvers. We compare our new solver with the existing state-of-the-art solvers and show its usefulness on synthetic and real datasets.
Zuzana Kukelova, Martin Bujnak, Tomás Pajdla
ICCV3
2012 Hand-Eye Calibration without Hand Orientation Measurement Using Minimal Solution
Zuzana Kukelova, Jan Heller, Tomás Pajdla
ACCV (4)3
2012 Making minimal solvers fast
abstract
In this paper we propose methods for speeding up minimal solvers based on Gröbner bases and action matrix eigenvalue computations. Almost all existing Gröbner basis solvers spend most time in the eigenvalue computation. We present two methods which speed up this phase of Gröbner basis solvers: (1) a method based on a modified FGLM algorithm for transforming Gröbner bases which results in a single-variable polynomial followed by direct calculation of its roots using Sturm-sequences and, for larger problems, (2) fast calculation of the characteristic polynomial of an action matrix, again solved using Sturm-sequences. We enhanced the FGLM method by replacing time consuming polynomial division performed in standard FGLM algorithm with efficient matrix-vector multiplication and we show how this method is related to the characteristic polynomial method. Our approaches allow computing roots only in some feasible interval and in desired precision. Proposed methods can significantly speedup many existing solvers. We demonstrate them on three important minimal computer vision problems.
Martin Bujnak, Zuzana Kukelova, Tomás Pajdla
CVPR3
2012 A branch-and-bound algorithm for globally optimal hand-eye calibration
abstract
This paper introduces a novel solution to hand-eye calibration problem. It is the first method that uses camera measurements directly and at the same time requires neither prior knowledge of the external camera calibrations nor a known calibration device. Our algorithm uses branch-and-bound approach to minimize an objective function based on the epipolar constraint. Further, it employs Linear Programming to decide the bounding step of the algorithm. The presented technique is able to recover both the unknown rotation and translation simultaneously and the solution is guaranteed to be globally optimal with respect to the L∞-norm.
Jan Heller, Michal Havlena, Tomás Pajdla
CVPR3
2012 Globally optimal hand-eye calibration
abstract
This paper introduces simultaneous globally optimal hand-eye self-calibration in both its rotational and translational components. The main contributions are new feasibility tests to integrate the hand-eye calibration problem into a branch-and-bound parameter space search. The presented method constitutes the first guaranteed globally optimal estimator for simultaneous optimization of both components with respect to a cost function based on reprojection errors. The system is evaluated in both synthetic and real world scenarios. The employed benchmark dataset is published online1to create a common point of reference for evaluation of hand-eye self-calibration algorithms.
Thomas Ruland, Tomás Pajdla
CVPR2
2012 Polynomial Eigenvalue Solutions to Minimal Problems in Computer Vision
abstract
We present a method for solving systems of polynomial equations appearing in computer vision. This method is based on polynomial eigenvalue solvers and is more straightforward and easier to implement than the state-of-the-art Gröbner basis method since eigenvalue problems are well studied, easy to understand, and efficient and robust algorithms for solving these problems are available. We provide a characterization of problems that can be efficiently solved as polynomial eigenvalue problems (PEPs) and present a resultant-based method for transforming a system of polynomial equations to a polynomial eigenvalue problem. We propose techniques that can be used to reduce the size of the computed polynomial eigenvalue problems. To show the applicability of the proposed polynomial eigenvalue method, we present the polynomial eigenvalue solutions to several important minimal relative pose problems.
Zuzana Kukelova, Martin Bujnak, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Beyond Novelty Detection: Incongruent Events, When General and Specific Classifiers Disagree
abstract
Unexpected stimuli are a challenge to any machine learning algorithm. Here, we identify distinct types of unexpected events when general-level and specific-level classifiers give conflicting predictions. We define a formal framework for the representation and processing of incongruent events: Starting from the notion of label hierarchy, we show how partial order on labels can be deduced from such hierarchies. For each event, we compute its probability in different ways, based on adjacent levels in the label hierarchy. An incongruent event is an event where the probability computed based on some more specific level is much smaller than the probability computed based on some more general level, leading to conflicting predictions. Algorithms are derived to detect incongruent events from different types of hierarchies, different applications, and a variety of data types. We present promising results for the detection of novel visual and audio objects, and new patterns of motion in video. We also discuss the detection of Out-Of- Vocabulary words in speech recognition, and the detection of incongruent events in a multimodal audiovisual scenario.
Daphna Weinshall, Alon Zweig, Hynek Hermansky, Stefan Kombrink, Frank W. Ohl, Jörn Anemüller, Jörg-Hendrik Bach, Luc Van Gool, Fabian Nater, Tomás Pajdla, Michal Havlena, Misha Pavel
IEEE Trans. Pattern Anal. Mach. Intell.10
2011 Structure-from-motion based hand-eye calibration using L∞ minimization
abstract
This paper presents a novel method for so-called hand-eye calibration. Using a calibration target is not possible for many applications of hand-eye calibration. In such situations Structure-from-Motion approach of hand-eye calibration is commonly used to recover the camera poses up to scaling. The presented method takes advantage of recent results in the L∞-norm optimization using Second-Order Cone Programming (SOCP) to recover the correct scale. Further, the correctly scaled displacement of the hand-eye transformation is recovered solely from the image correspondences and robot measurements, and is guaranteed to be globally optimal with respect to the L∞-norm. The method is experimentally validated using both synthetic and real world datasets.
Jan Heller, Michal Havlena, Akihiro Sugimoto, Tomás Pajdla
CVPR4
2011 Multi-view reconstruction preserving weakly-supported surfaces
abstract
We propose a novel method for the multi-view reconstruction problem. Surfaces which do not have direct support in the input 3D point cloud and hence need not be photo-consistent but represent real parts of the scene (e.g. low-textured walls, windows, cars) are important for achieving complete reconstructions. We augmented the existing Labatut CGF 2009 method with the ability to cope with these difficult surfaces just by changing the t-edge weights in the construction of surfaces by a minimal s-t cut. Our method uses Visual-Hull to reconstruct the difficult surfaces which are not sampled densely enough by the input 3D point cloud. We demonstrate importance of these surfaces on several real-world data sets. We compare our improvement to our implementation of the Labatut CGF 2009 method and show that our method can considerably better reconstruct difficult surfaces while preserving thin structures and details in the same quality and computational time.
Michal Jancosek, Tomás Pajdla
CVPR2
2011 Global optimization of extended hand-eye calibration
abstract
This paper introduces simultaneous global optimization of both camera orientation and vehicle wheel circumference without requiring any information about the translations in the system. The main contribution are new objective function bounds to integrate this problem into a branch-and-bound parameter space search. The presented method constitutes the first guaranteed globally optimal estimator for both components of the problem with respect to a cost function based on reprojection errors. The algorithm operates directly on image measurements and does not depend on any structure and motion preprocessing to estimate camera poses. The complete system is implemented and validated on both synthetic and real automotive datasets.
Thomas Ruland, Tomás Pajdla
Intelligent Vehicles Symposium2
2011 Omnidirectional Image Stabilization for Visual Object Recognition
Akihiko Torii, Michal Havlena, Tomás Pajdla
Int. J. Comput. Vis.3
2011 A Minimal Solution to Radial Distortion Autocalibration
abstract
Simultaneous estimation of radial distortion, epipolar geometry, and relative camera pose can be formulated as a minimal problem and solved from a minimal number of image points. Finding the solution to this problem leads to solving a system of algebraic equations. In this paper, we provide two different solutions to the problem of estimating radial distortion and epipolar geometry from eight point correspondences in two images. Unlike previous algorithms which were able to solve the problem from nine correspondences only, we enforce the determinant of the fundamental matrix be zero. This leads to a system of eight quadratic and one cubic equation in nine variables. We first simplify this system by eliminating six of these variables and then solve the system by two alternative techniques. The first one is based on the Gröbner basis method and the second one on the polynomial eigenvalue computation. We demonstrate that our solutions are efficient, robust, and practical by experiments on synthetic and real data.
Zuzana Kukelova, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 New Efficient Solution to the Absolute Pose Problem for Camera with Unknown Focal Length and Radial Distortion
Martin Bujnak, Zuzana Kukelova, Tomás Pajdla
ACCV (1)3
2010 Closed-Form Solutions to Minimal Absolute Pose Problems with Known Vertical Direction
Zuzana Kukelova, Martin Bujnak, Tomás Pajdla
ACCV (2)3
2010 Simultaneous surveillance camera calibration and foot-head homology estimation from human detections
abstract
We propose a novel method for automatic camera calibration and foot-head homology estimation by observing persons standing at several positions in the camera field of view. We demonstrate that human body can be considered as a calibration target thus avoiding special calibration objects or manually established fiducial points. First, by assuming roughly parallel human poses we derive a new constraint which allows to formulate the calibration of internal and external camera parameters as a Quadratic Eigenvalue Problem. Secondly, we couple the calibration with an improved effective integral contour based human detector and use 3D projected models to capture a large variety of person and camera mutual positions. The resulting camera auto-calibration method is very robust and efficient, and thus well suited for surveillance applications where the camera calibration process cannot use special calibration targets and must be simple.
Branislav Micusík, Tomás Pajdla
CVPR2
2010 Efficient Structure from Motion by Graph Optimization
Michal Havlena, Akihiko Torii, Tomás Pajdla
ECCV (2)3
2010 Avoiding Confusing Features in Place Recognition
Jan Knopp, Josef Sivic, Tomás Pajdla
ECCV (1)3
2010 Special issue on omnidirectional vision, camera networks and non-conventional cameras
João Pedro Barreto 0001, Tomás Pajdla, Akihiro Sugimoto
Comput. Vis. Image Underst.2
2010 Fast and robust numerical solutions to minimal problems for cameras with radial distortion
Zuzana Kukelova, Martin Byröd, Klas Josephson, Tomás Pajdla, Kalle Åström
Comput. Vis. Image Underst.4
2009 Robust Focal Length Estimation by Voting in Multi-view Scene Reconstruction
Martin Bujnak, Zuzana Kukelova, Tomás Pajdla
ACCV (1)3
2009 Randomized structure from motion based on atomic 3D models from camera triplets
abstract
This paper presents a new efficient technique for large-scale structure from motion from unordered data sets. We avoid costly computation of all pairwise matches and geometries by sampling pairs of images using the pairwise similarity scores based on the detected occurrences of visual words leading to a significant speedup. Furthermore, atomic 3D models reconstructed from camera triplets are used as the seeds which form the final large-scale 3D model when merged together. Using three views instead of two allows us to reveal most of the outliers of pairwise geometries at an early stage of the process hindering them from derogating the quality of the resulting 3D structure at later stages. The accuracy of the proposed technique is shown on a set of 64 images where the result of the exhaustive technique is known. Scalability is demonstrated on a landmark reconstruction from hundreds of images.
Michal Havlena, Akihiko Torii, Jan Knopp, Tomás Pajdla
CVPR4
2009 Stereographic rectification of omnidirectional stereo pairs
abstract
We present a general technique for rectification of a stereo pair acquired by a calibrated omnidirectional camera. Using this technique we formulate a new stereographic rectification method. Our rectification does not map epipolar curves onto lines as common rectification methods, but rather maps epipolar curves onto circles. We show that this rectification in a certain sense minimizes the distortion of the original omnidirectional images. We formulate the rectification for multiple images and show that the choice of the optimal projection center of the rectification is under certain circumstances equivalent to the classical problem of spherical minimax location. We demonstrate the behaviour and the quality of the rectification in real experiments with images from 180 degree field of view fish eye lenses.
Jan Heller, Tomás Pajdla
CVPR2
2009 3D reconstruction from image collections with a single known focal length
abstract
In this paper we aim at reconstructing 3D scenes from images with unknown focal lengths downloaded from photosharing websites such as Flickr. First we provide a minimal solution to finding the relative pose between a completely calibrated camera and a camera with an unknown focal length given six point correspondences. We show that this problem has up to nine solutions in general and present two efficient solvers to the problem. They are based on Gröbner basis, resp. on generalized eigenvalues, computation. We demonstrate by experiments with synthetic and real data that both solvers are correct, fast, numerically stable and work well even in some situations when the classical 6-point algorithm fails, e.g. when optical axes of the cameras are parallel or intersecting. Based on this solution we present a new efficient method for large-scale structure from motion from unordered data sets downloaded from the Internet. We show that this method can be effectively used to reconstruct 3D scenes from collection of images with very few (in principle single) images with known focal lengths.
Martin Bujnak, Zuzana Kukelova, Tomás Pajdla
ICCV3
2009 Omnidirectional Image Stabilization by Computing Camera Trajectory
Akihiko Torii, Michal Havlena, Tomás Pajdla
PSIVT3
2008 Polynomial Eigenvalue Solutions to the 5-pt and 6-pt Relative Pose Problems
abstract
In this paper we provide new fast and simple solutions to two important minimal problems in computer vision, the five-point relative pose problem and the six-point focal length problem. We show that these two problems can easily be formulated as polynomial eigenvalue problems of degree three and two and solved using standard efficient numerical algorithms. Our solutions are somewhat more stable than state-of-the-art solutions by Nister and Stewenius and are in some sense more straightforward and easier to implement since polynomial eigenvalue problems are well studied with many efficient and robust algorithms available. The quality of the solvers is demonstrated in experiments 1. 1
Zuzana Kukelova, Martin Bujnak, Tomás Pajdla
BMVC3
2008 A general solution to the P4P problem for camera with unknown focal length
abstract
This paper presents a general solution to the determination of the pose of a perspective camera with unknown focal length from images of four 3D reference points. Our problem is a generalization of the P3P and P4P problems previously developed for fully calibrated cameras. Given four 2D-to-3D correspondences, we estimate camera position, orientation and recover the camera focal length. We formulate the problem and provide a minimal solution from four points by solving a system of algebraic equations. We compare the Hidden variable resultant and Grobner basis techniques for solving the algebraic equations of our problem. By evaluating them on synthetic and on real-data, we show that the Grobner basis technique provides stable results.
Martin Bujnak, Zuzana Kukelova, Tomás Pajdla
CVPR3
2008 Fast and robust numerical solutions to minimal problems for cameras with radial distortion
abstract
A number of minimal problems of structure from motion for cameras with radial distortion have recently been studied and solved in some cases. These problems are known to be numerically very challenging and in several cases there exist no known practical algorithm yielding solutions in floating point arithmetic. We make some crucial observations concerning the floating point implementation of Gröbner basis computations and use these new insights to formulate fast and stable algorithms for two minimal problems with radial distortion previously solved in exact rational arithmetic only: (i) simultaneous estimation of essential matrix and a common radial distortion parameter for two partially calibrated views and six image point correspondences and (ii) estimation of fundamental matrix and two different radial distortion parameters for two uncalibrated views and nine image point correspondences. We demonstrate on simulated and real experiments that these two problems can be efficiently solved in floating point arithmetic.
Martin Byröd, Zuzana Kukelova, Klas Josephson, Tomás Pajdla, Kalle Åström
CVPR4
2008 Measuring camera translation by the dominant apical angle
abstract
This paper provides a technique for measuring camera translation relatively w.r.t. the scene from two images. We demonstrate that the amount of the translation can be reliably measured for general as well as planar scenes by the most frequent apical angle, the angle under which the camera centers are seen from the perspective of the reconstructed scene points. Simulated experiments show that the dominant apical angle is a linear function of the length of the true camera translation. In a real experiment, we demonstrate that by skipping image pairs with too small motion, we can reliably initialize structure from motion, compute accurate camera trajectory in order to rectify images and use the ground plane constraint in recognition of pedestrians in a hand-held video sequence.
Akihiko Torii, Michal Havlena, Tomás Pajdla, Bastian Leibe
CVPR3
2008 Automatic Generator of Minimal Problem Solvers
Zuzana Kukelova, Martin Bujnak, Tomás Pajdla
ECCV (3)3
2008 The DIRAC AWEAR audio-visual platform for detection of unexpected and incongruent events
abstract
It is of prime importance in everyday human life to cope with and respond appropriately to events that are not foreseen by prior experience. Machines to a large extent lack the ability to respond appropriately to such inputs. An important class of unexpected events is defined by incongruent combinations of inputs from different modalities and therefore multimodal information provides a crucial cue for the identification of such events, e.g., the sound of a voice is being heard while the person in the field-of-view does not move her lips. In the project DIRAC ("Detection and Identification of Rare Audio-visual Cues") we have been developing algorithmic approaches to the detection of such events, as well as an experimental hardware platform to test it. An audio-visual platform ("AWEAR" - audio-visual wearable device) has been constructed with the goal to help users with disabilities or a high cognitive load to deal with unexpected events. Key hardware components include stereo panoramic vision sensors and 6-channel worn-behind-the-ear (hearing aid) microphone arrays. Data have been recorded to study audio-visual tracking, a/v scene/object classification and a/v detection of incongruencies.
Jörn Anemüller, Jörg-Hendrik Bach, Barbara Caputo, Michal Havlena, Jie Luo 0018, Hendrik Kayser, Bastian Leibe, Petr Motlícek, Tomás Pajdla, Misha Pavel, Akihiko Torii, Luc Van Gool, Alon Zweig, Hynek Hermansky
ICMI9
2008 Eliminating Blind Spots for Assisted Driving
abstract
Drivers of heavy goods vehicles are not able to survey the whole surrounding area of their vehicle due to large blind spot regions. This paper shows how catadioptric cameras—a combination of cameras and mirrors—can be used to survey the surrounding area of vehicles. Four such cameras were mounted on a truck–trailer combination, and the images are combined such that obstacles are visible in an image presented to the driver. This image is a bird's eye view of the vehicle. Additionally, corridors indicating the path of motion of the vehicle are overlaid to the resulting image. To compute those corridors, a mathematical description of the path of motion is derived. Such a system does not only support the driver during maneuvering tasks but also increases safety of driving large vehicles.
Tobias Ehlgen, Tomás Pajdla, D. Ammon
IEEE Trans. Intell. Transp. Syst.2
2007 A minimal solution to the autocalibration of radial distortion
abstract
Epipolar geometry and relative camera pose computation are examples of tasks which can be formulated as minimal problems and solved from a minimal number of image points. Finding the solution leads to solving systems of algebraic equations. Often, these systems are not trivial and therefore special algorithms have to be designed to achieve numerical robustness and computational efficiency. In this paper we provide a solution to the problem of estimating radial distortion and epipolar geometry from eight correspondences in two images. Unlike previous algorithms, which were able to solve the problem from nine correspondences only, we enforce the determinant of the fundamental matrix be zero. This leads to a system of eight quadratic and one cubic equation in nine variables. We simplify this system by eliminating six of these variables. Then, we solve the system by finding eigenvectors of an action matrix of a suitably chosen polynomial. We show how to construct the action matrix without computing complete Grobner basis, which provides an efficient and robust solver. The quality of the solver is demonstrated on synthetic and real data.
Zuzana Kukelova, Tomás Pajdla
CVPR2
2007 Robust Rotation and Translation Estimation in Multiview Reconstruction
abstract
It is known that the problem of multiview reconstruction can be solved in two steps: first estimate camera rotations and then translations using them. This paper presents new robust techniques for both of these steps. (i) Given pairwise relative rotations, global camera rotations are estimated linearly in least squares. (ii) Camera translations are estimated using a standard technique based on Second Order Cone Programming. Robustness is achieved by using only a subset of points according to a new criterion that diminishes the risk of chosing a mismatch. It is shown that only four points chosen in a special way are sufficient to represent a pairwise reconstruction almost equally as all points. This leads to a significant speedup. In image sets with repetitive or similar structures, non-existent epipolar geometries may be found. Due to them, some rotations and consequently translations may be estimated incorrectly. It is shown that iterative removal of pairwise reconstructions with the largest residual and reregistration removes most non-existent epipolar geometries. The performance of the proposed method is demonstrated on difficult wide base-line image sets.
Daniel Martinec, Tomás Pajdla
CVPR2
2007 Multi-label image segmentation via max-sum solver
abstract
We formulate single-image multi-label segmentation into regions coherent in texture and color as a MAX-SUM problem for which efficient linear programming based solvers have recently appeared. By handling more than two labels, we go beyond widespread binary segmentation methods, e.g., MIN-CUT or normalized cut based approaches. We show that the MAX-SUM solver is a very powerful tool for obtaining the MAP estimate of a Markov random field (MRF). We build the MRF on superpixels to speed up the segmentation while preserving color and texture. We propose new quality functions for setting the MRF, exploiting priors from small representative image seeds, provided either manually or automatically. We show that the proposed automatic segmentation method outperforms previous techniques in terms of the global consistency error evaluated on the Berkeley segmentation database.
Branislav Micusík, Tomás Pajdla
CVPR2
2007 Two Minimal Problems for Cameras with Radial Distortion
abstract
Epipolar geometry and relative camera pose computation for uncalibrated cameras with radial distortion has recently been formulated as a minimal problem and successfully solved in floating point arithmetics. The singularity of the fundamental matrix has been used to reduce the minimal number of points to eight. It was assumed that the cameras were not calibrated but had same distortions. In this paper we formulate two new minimal problems for estimating epipolar geometry of cameras with radial distortion. First we present a minimal algorithm for partially calibrated cameras with same radial distortion. Using the trace constraint which holds for the epipolar geometry of calibrated cameras to reduce the number of necessary points from eight to six. We demonstrate that the problem is solvable in exact rational arithmetics. Secondly, we present a minimal algorithm for uncalibrated cameras with different radial distortions. We show that the problem can be solved using nine points in two views by manipulating polynomials by a sequence of Gauss-Jordan eliminations in exact rational arithmetics. We demonstrate the algorithms on synthetic and real data.
Zuzana Kukelova, Tomás Pajdla
ICCV2
2007 Maneuvering Aid for Large Vehicle using Omnidirectional Cameras
abstract
Maneuvering large vehicles like trucks is a challenging task because their drivers often cannot see the closest surrounding area of the vehicle. This paper presents a system providing the driver with a bird's-eye view of the surrounding area on a single display. This view is constructed by stitching together cutouts from two omnidirectional images acquired by catadioptric cameras, which are placed below the side mirrors. The geometry of the cameras is analyzed with respect to the formation of blind spots. We illustrate how to select the cutouts in order to eliminate completely the blind spots in the bird's-eye view
Tobias Ehlgen, Tomás Pajdla
WACV2
2006 Editorial: Selection of Papers for the ECCV 2004 Special Issue
Tomás Pajdla, Jiri Matas
Int. J. Comput. Vis.1
2006 Structure from Motion with Wide Circular Field of View Cameras
abstract
This paper presents a method for fully automatic and robust estimation of two-view geometry, autocalibration, and 3D metric reconstruction from point correspondences in images taken by cameras with wide circular field of view. We focus on cameras which have more than 180 degrees field of view and for which the standard perspective camera model is not sufficient, e.g., the cameras equipped with circular fish-eye lenses Nikon FC-E8 (183 degrees), Sigma 8mm-f4-EX (180 degrees), or with curved conical mirrors. We assume a circular field of view and axially symmetric image projection to autocalibrate the cameras. Many wide field of view cameras can still be modeled by the central projection followed by a nonlinear image mapping. Examples are the above-mentioned fish-eye lenses and properly assembled catadioptric cameras with conical mirrors. We show that epipolar geometry of these cameras can be estimated from a small number of correspondences by solving a polynomial eigenvalue problem. This allows the use of efficient RANSAC robust estimation to find the image projection model, the epipolar geometry, and the selection of true point correspondences from tentative correspondences contaminated by mismatches. Real catadioptric cameras are often slightly noncentral. We show that the proposed autocalibration with approximate central models is usually good enough to get correct point correspondences which can be used with accurate noncentral models in a bundle adjustment to obtain accurate 3D scene reconstruction. Noncentral camera models are dealt with and results are shown for catadioptric cameras with parabolic and spherical mirrors.
Branislav Micusík, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 3D Reconstruction by Fitting Low-Rank Matrices with Missing Data
abstract
A technique for building consistent 3D reconstructions from many views based on fitting a low rank matrix to a matrix with missing data is presented. Rank-four submatrices of minimal, or slightly larger, size are sampled and spans of their columns are combined to constrain a basis of the fitted matrix. The error minimized is expressed in terms of the original subspaces which leads to a better resistance to noise compared to previous methods. More than 90% of the missing data can be handled while finding an acceptable solution efficiently. Applications to 3D reconstruction using both affine and perspective camera models are shown. For the perspective model, a new linear method based on logarithms of positive depths from chirality is introduced to make the depths consistent with an overdetermined set of epipolar geometries. Results are shown for scenes and sequences of various types. Many images in open and closed sequences in narrow and wide base-line setups are reconstructed with reprojection errors around one pixel. It is shown that reconstructed cameras can be used to obtain dense reconstructions from epipolarly aligned images.
Daniel Martinec, Tomás Pajdla
CVPR (1)2
2005 The geometric error for homographies
Ondrej Chum, Tomás Pajdla, Peter F. Sturm
Comput. Vis. Image Underst.2
2004 Constraints on perspective images and circular panoramas
abstract
We describe an algebraic constraint on corresponding image points in a perspective image and a circular panorama and provide a method to estimate it from noisy image measurements. Studying this combination of cameras is a step forward in localization and recognition since a database of circular panoramas captures completely the appearance of objects and scenes, and perspective images are the simplest query images. The constraint gives a way to use a RANSAC-like algorithm for image matching. We introduce a general method to establish constraints between (non-central) images in the form of a bilinear function of the lifted coordinates of corresponding image points. We apply the method to obtain an algebraic constraint for a perspective image and a circular panorama. The algebraic constraints are interpreted geometrically and the constraints estimated from image data are used to auto-calibrate cameras and to compute a metric reconstruction of the scene observed. A synthetic experiment demonstrates that the proposed reconstruction method behaves favorably in presence of image noise. As a proof of concept, the constraints are estimated from real images of indoor scenes and used to reconstruct positions of cameras and to compute a metric reconstruction of the scene.
Marc Menem, Tomás Pajdla
BMVC2
2004 Autocalibration & 3D Reconstruction with Non-Central Catadioptric Cameras
Branislav Micusík, Tomás Pajdla
CVPR (1)2
2004 Robust wide-baseline stereo from maximally stable extremal regions
Jiri Matas, Ondrej Chum, Tomás Pajdla
Image Vis. Comput.4
2003 Rendering Almost Perspective Views from a Sparse Set of Omnidirectional Images
abstract
Non-central (X-slits) images can be used in image based rendering (IBR) for new view synthesis for a viewer moving in a restricted region. Novel, i.e. previously unseen, views are rendered from virtual camera positions from a set of images captured at some predefined camera positions. The novel views correspond to camera positions not contained in the input image set. For instance, the input sequence consists of images captured on a circular path, but the viewer can move in a disk defined by that path. We investigate a class of non-central images, the omnidirectional X-slits images, which capture environment around the viewer completely. The virtual viewer can rotate his head freely in each position without the need of creating another image for different viewing direction. New images have to be generated only in case of a change in position of the viewer. This approach allows efficient representation of high resolution (and thus space demanding) images for IBR in memory. 1
Hynek Bakstein, Tomás Pajdla, Daniel Vecerka
BMVC2
2003 Joint Orientation of Epipoles
abstract
It is known that epipolar constraint can be augmented with orientation by formulating it in the oriented projective geometry. This oriented epipolar constraint requires knowing the orientations (signs of overall scales) of epipoles and fundamental matrix. The current belief is that these orientations cannot be obtained from the fundamental matrix only and that additional information is needed, typically, a single correct point correspondence. In contrary to this, we show that fundamental matrix alone encodes orientation of epipoles up to their common scale sign. We present two formulations of this fact. The algebraic formulation gives a closed formula to compute the second epipole from fundamental matrix and the first epipole. The geometric formulation is in terms of the conic formed by intersections of corresponding epipolar lines in the common image plane; we show that the epipoles always lie on different antipodal components of the spherical interpretation of this conic. Further, we show that, under mild assumptions, fundamental matrix can discriminate between two classes of mutual position of a pair of directional cameras. 1
Ondrej Chum, Tomás Werner, Tomás Pajdla
BMVC3
2003 Line Reconstruction from Many Perspective Images by Factorization
abstract
This paper proposes a method for line reconstruction from many perspective images by factorization of a matrix containing line correspondences. No point correspondences are used. We formulate the reconstruction from line correspondences in the language of Plucker line coordinates. The reconstruction is posed as the factorization of 3m /spl times/ n matrix S into the product S = QL of 3m /spl times/ 6 projection matrix Q and 6 /spl times/ n line matrix L, both satisfying Klein identities. The matrix S contains coordinates of lines detected in perspective images. Similarly to reconstruction from point correspondences in perspective images, the matrix S has to be properly rescaled before it can be factorized. We propose a scaling of image line coordinates based on trifocal tensors that are analogical to the scaling proposed by Sturm and Triggs (1996) for points. We propose an SVD based factorization enforcing Klein identities on Q and L in a noise-free situation. We show experiments on real data that suggest that a good reconstruction may be obtained even if data is noisy and the identities are not enforced exactly. We also discuss an extension of the method for images with occlusions.
Daniel Martinec, Tomás Pajdla
CVPR (1)2
2003 Estimation of omnidirectional camera model from epipolar geometry
abstract
We generalize the method of simultaneous linear estimation of multiple view geometry and lens distortion, introduced by Fitzgibbon at CVPR 2001, to an omnidirectional (angle of view larger than 180/spl deg/) camera. The perspective camera is replaced by a linear camera with a spherical retina and a nonlinear mapping of the sphere into the image plane. Unlike the previous distortion-based models, the new camera model is capable to describe a camera with an angle of view larger than 180/spl deg/ at the cost of introducing only one extra parameter. A suitable linearization of the camera model and of the epipolar constraint is developed in order to arrive at a quadratic eigenvalue problem for which efficient algorithms are known. The lens calibration is done from automatically established image correspondences only. Besides rigidity, no assumptions about the scene are made (e.g. presence of a calibration object). We demonstrate the method in experiments with Nikon FC-E8 fish-eye converter for COOLPIX digital camera. In practical situations, the proposed method allows to incorporate the new omnidirectional camera model into RANSAC - a robust estimation technique.
Branislav Micusík, Tomás Pajdla
CVPR (1)2
2003 On the Epipolar Geometry of the Crossed-Slits Projection
abstract
The Crossed-Slits (X-Slits) camera is defined by two nonintersecting slits, which replace the pinhole in the common perspective camera. Each point in space is projected to the image plane by a ray which passes through the point and the two slits. The X-Slits projection model includes the pushbroom camera as a special case. In addition, it describes a certain class of panoramic images, which are generated from sequences obtained by translating pinhole cameras. In this paper we develop the epipolar geometry of the X-Slits projection model. We show an object which is similar to the fundamental matrix; our matrix, however, describes a quadratic relation between corresponding image points (using the Veronese mapping). Similarly the equivalent of epipolar lines are conics in the image plane. Unlike the pin-hole case, epipolar surfaces do not usually exist in the sense that matching epipolar lines lie on a single surface; we analyze the cases when epipolar surfaces exist, and characterize their properties. Finally, we demonstrate the matching of points in pairs of X-Slits panoramic images.
Doron Feldman, Tomás Pajdla, Daphna Weinshall
ICCV2
2002 Robust Wide Baseline Stereo from Maximally Stable Extremal Regions
abstract
Abstract The wide-baseline stereo problem, i.e. the problem of establishing correspondences between a pair of images taken from different viewpoints is studied. A new set of image elements that are put into correspondence, the so called extremal regions , is introduced. Extremal regions possess highly desirable properties: the set is closed under (1) continuous (and thus projective) transformation of image coordinates and (2) monotonic transformation of image intensities. An efficient (near linear complexity) and practically fast detection algorithm (near frame rate) is presented for an affinely invariant stable subset of extremal regions, the maximally stable extremal regions (MSER). A new robust similarity measure for establishing tentative correspondences is proposed. The robustness ensures that invariants from multiple measurement regions (regions obtained by invariant constructions from extremal regions), some that are significantly larger (and hence discriminative) than the MSERs, may be used to establish tentative correspondences. The high utility of MSERs, multiple measurement regions and the robust metric is demonstrated in wide-baseline experiments on image pairs from both indoor and outdoor scenes. Significant change of scale (3.5×), illumination conditions, out-of-plane rotation, occlusion, locally anisotropic scale change and 3D translation of the viewpoint are all present in the test problems. Good estimates of epipolar geometry (average distance from corresponding points to the epipolar line below 0.09 of the inter-pixel distance) are obtained.
Jiri Matas, Ondrej Chum, Tomás Pajdla
BMVC4
2002 Structure from Many Perspective Images with Occlusions
Daniel Martinec, Tomás Pajdla
ECCV (2)2
2002 Differential Invariants as the Base of Triangulated Surface Registration
Pavel Krsek, Tomás Pajdla, Václav Hlavác
Comput. Vis. Image Underst.2
2002 Stereo with Oblique Cameras
Tomás Pajdla
Int. J. Comput. Vis.1
2002 Epipolar Geometry for Central Catadioptric Cameras
Tomás Svoboda, Tomás Pajdla
Int. J. Comput. Vis.2
2001 Oriented Matching Constraints
abstract
Well-known matching constraints for points and lines in muliple images are necessary but not sufficient condition for the existence of real structure and cameras, underlying the image correspondences. To obtain sufficient conditions, the following additional constraints must be imposed: positive scales, the existence of a plane at infinity not intersecting the scene, and the existence of handedness preserving cameras. We present modifications of the well-known matching constraints and also some new constraints, taking into account some of this additional knowledge. Not only conventional but also central panoramic cameras are naturally described. To achieve this, we have generalized and simplified Hartley's ch(e)irality theory by formulating it in the language of oriented projective geometry and Grassmann tensors.
Tomás Werner, Tomás Pajdla
BMVC2
2001 Matching in Catadioptric Images with Appropriate Windows, and Outliers Removal
Tomás Svoboda, Tomás Pajdla
CAIP2
2001 3D Reconstruction from 360 x 360 Mosaics
abstract
We are studying the geometry of a 360/spl times/360 mosaic image formation. A 360/spl times/360 mosaic camera model and a calibration procedure are proposed. It is shown that only one point correspondence is needed in order to acquire epipolar rectified images. The 360/spl times/360 mosaic camera model is therefore determined by only one intrinsic parameter It is shown that the relation between coordinates estimated with different values of intrinsic 360/spl times/360 mosaic camera parameters is a scaling of all scene point coordinates with additional nonlinear changes in the z coordinates of the scene points. Experimental results verifying the reconstruction of real scene points are presented.
Hynek Bakstein, Tomás Pajdla
CVPR (2)2
2001 Cheirality in Epipolar Geometry
Tomás Werner, Tomás Pajdla
ICCV2
1999 Zero Phase Representation of Panoramic Images for Image Vased Localization
Tomás Pajdla, Václav Hlavác
CAIP1
1998 Camera Calibration and Euclidean Reconstruction from Known Observer Translations
abstract
We present a technique for camera calibration and Euclidean reconstruction from multiple images of the same scene. Unlike standard Tsai's camera calibration from a known scene, we exploited controlled known motions of the camera to obtain its calibration and Euclidean reconstruction without any knowledge about the scene. We consider three linearly independent translations of an uncalibrated camera mounted on a robot arm that provides us with four views of the scene. The translations of the robot arm are measured in a robot coordinate system. This special, but still realistic, arrangement allowed us to find a linear algorithm for recovering all intrinsic camera calibration parameters, the rotation of the camera with respect to the robot coordinate system, and proper scaling factors for all points allowing their Euclidean reconstruction. The experiments showed that an efficient and robust algorithm was obtained by exploiting Total Least Squares in combination with careful normalization of image coordinates.
Tomás Pajdla, Václav Hlavác
CVPR1
1998 Epipolar Geometry of Panoramic Cameras
Tomás Svoboda, Tomás Pajdla, Václav Hlavác
ECCV (1)2
1998 Efficient 3-D Scene Visualization by Image Extrapolation
Tomás Werner, Tomás Pajdla, Václav Hlavác
ECCV (2)2
1998 Efficient rendering of projective model for image-based visualization
abstract
This work describes a method for synthesizing correct virtual images of a real scene, described by a set of uncalibrated reference images. The approach used is rendering a projective 3D model. Projective model reconstruction and positioning a projective virtual camera are not addressed. First, the algorithm transfers the vertices of triangles in the triangulated model and then warps interiors of triangles. Hidden faces are removed by z-buffering. The algorithm is more efficient than the ray-tracing-like algorithm for virtual view synthesis by Laveau and Faugeras (1996).
Tomás Werner, Tomás Pajdla, Václav Hlavác
ICPR2
1996 Selection of reference views for image-based representation
abstract
Recently, much attention has been devoted to image-based scene representations. They allow one to construct an arbitrary view of a 3D scene by the interpolation (transfer) from a sparse set of real 2D (reference) images, rather than by rendering an explicit 3D model. While many authors address mainly the purely geometrical aspect of the task, we focus on the problem of how to select the optimal set of reference views. Selection of reference views from a dense set of real primary views is posed as a selection and fitting of parametric models. The selected set must minimize a weighted sum of the number of reference views and the total fit error. We propose two different algorithms for solving this optimization problem. The experimental results on synthetic and real data indicate the feasibility of the approach for 1-DOF camera movement. We discuss the possibility to extend one of the algorithms for more general case.
Tomás Werner, Václav Hlavác, Ales Leonardis, Tomás Pajdla
ICPR4
1996 Choosing Reference Views for Image-Based Representation
Tomás Werner, Václav Hlavác, Ales Leonardis, Tomás Pajdla
SOFSEM4
1995 Efficient matching of space curves
Tomás Pajdla, Luc Van Gool
CAIP1
1995 Matching of 3D Curves Using Semi-Differential Invariants
abstract
A method for matching 3-D curves under Euclidean motions is presented. Our approach uses a semi-differential invariant description requiring only first derivatives and one reference point, thus avoiding the computation of high order derivatives. A novel curve similarity measure building on the notion of /spl epsiv/-reciprocal correspondence is proposed. It is shown that by combining /spl epsiv/-reciprocal correspondence with the robust least median of squares motion estimation, the registration of partially occluded curves can be accomplished. An experiment with real curves extracted from 3-D surfaces demonstrates that curve matching can be successfully performed even on data from a simple and cheap 3-D sensor.>
Tomás Pajdla, Luc Van Gool
ICCV1
1994 Improvement of the curvature computation
abstract
The improvement of computing the the curvature of the digitized curves is presented. The standard scheme, i.e. computing curvature using convolution with the truncated Gaussian kernel, was studied. First, we show that systematic bias caused by curvature smoothing can be removed. Second, we demonstrate that large portion of the error has roots in other phenomena (i.e. anisotropy of the raster, limited size of the Gaussian, numerical integration of the convolution, and discretization).
Václav Hlavác, Tomás Pajdla, Milos Sommer
ICPR (1)2
1993 Surface Discontinuities in Range Images
Tomás Pajdla, Václav Hlavác
CAIP1
1993 Surface discontinuities in range images
abstract
The authors present a formulation of discontinuity detection in range images. The emphasis is on roof edge detection which is evaluated as C/sup 1/ discontinuity strength at every point in the range image. Contrary to other local approaches, it is possible to detect all shapes of C/sup 1/ discontinuity from one polynomial approximation even when the viewpoint changes. First, the range data are locally approximated by a second-order bivariate polynomial. Second, the 2-D problem of surface discontinuity strength estimation is converted to a 1-D probem along the direction of maximal normal curvature at each point on the surface. Third, C/sup 1/ is computed by comparing the polynomial approximation of data and the polynomial approximation of the discretized model of the discontinuity.>
Tomás Pajdla, Václav Hlavác
ICCV1