EDBT 2026 Demo / reviewers in the wild / expert
Daniel Barath
dblp:166/4789 · also Dániel Baráth, Dániel Béla Baráth
· DBLP profile ↗
82ranked-venue papers
30as first author
59since 2021 · last 2025
0000-0002-8736-0222ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 27 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 68 · 25 first-author · 49 since 2021Systems, architecture and hardware · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Filter Outlier Edges in Global SfMabstractWe present a novel approach to enhance camera pose estimation in global Structure-from-Motion (SfM) frameworks by filtering inaccurate pose graph edges – representing relative translation estimates – before applying translation averaging. In SfM, pose graph vertices represent images, and edges represent relative poses (rotations and translations) between cameras. We reformulate the edge filtering problem as a vertex filtering in the dual graph, specifically, a line graph where vertices correspond to edges in the original graph and edges correspond to cameras. Utilizing this representation, we frame the problem as a binary classification over nodes in the dual graph. To identify outlier edges, we employ a Transformer-based architecture. To overcome the challenge of memory overflow caused by converting to a line graph, we introduce a clustering-based graph processing approach, enabling our method to be applied to arbitrarily large pose graphs. Our method outperforms existing relative translation filtering techniques in terms of camera position accuracy and can be seamlessly integrated with other filters. The code is available at https://github.com/DmblnNicole/LFOE-GlobalSfM. Nicole Damblon, Marc Pollefeys, Daniel Barath |
CVPR | 3 |
| 2025 | CrossOver: 3D Scene Cross-Modal AlignmentabstractMulti-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment. Unlike traditional methods that require aligned modality data for every object instance, CrossOver learns a unified, modality-agnostic embedding space for scenes by aligning modalities – RGB images, point clouds, CAD models, floorplans, and text descriptions – with relaxed constraints and without explicit object semantics. Leveraging dimensionality-specific encoders, a multi-stage training pipeline, and emergent cross-modal behaviors, CrossOver supports robust scene retrieval and object localization, even with missing modalities. Evaluations on Scan-Net and 3RScan datasets show its superior performance across diverse metrics, highlighting CrossOver’s adaptability for real-world applications in 3D scene understanding. Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni |
CVPR | 4 |
| 2025 | Learning Affine Correspondences by Integrating Geometric ConstraintsabstractAffine correspondences have received significant attention due to their benefits in tasks like image matching and pose estimation. Existing methods for extracting affine correspondences still have many limitations in terms of performance; thus, exploring a new paradigm is crucial. In this paper, we present a new pipeline designed for extracting accurate affine correspondences by integrating dense matching and geometric constraints. Specifically, a novel extraction framework is introduced, with the aid of dense matching and a novel keypoint scale and orientation estimator. For this purpose, we propose loss functions based on geometric constraints, which can effectively improve accuracy by supervising neural networks to learn feature geometry. The experimental show that the accuracy and robustness of our method outperform the existing ones in image matching tasks. To further demonstrate the effectiveness of the proposed method, we applied it to relative pose estimation. Affine correspondences extracted by our method lead to more accurate poses than the baselines on a range of real-world datasets. The code is available at https://github.com/stilcrad/LearningACs. Pengju Sun, Banglei Guan, Zhenbao Yu, Yang Shang, Daniel Barath |
CVPR | 6 |
| 2025 | Practical Solutions to the Relative Pose of Three Calibrated CamerasabstractWe study the challenging problem of estimating the relative pose of three calibrated cameras from four point correspondences. We propose novel efficient solutions to this problem that are based on the simple idea of using four correspondences to estimate an approximate geometry of the first two views. We model this geometry either as an affine or a fully perspective geometry estimated using one additional approximate correspondence. We generate such an approximate correspondence using a very simple and efficient strategy, where the new point is the mean point of three corresponding input points. The new solvers are efficient and easy to implement, since they are based on existing efficient minimal solvers, i.e., the 4-point affine fundamental matrix, the well-known 5-point relative pose solver, and the P3P solver. Extensive experiments on real data show that the proposed solvers, when properly coupled with local optimization, achieve state-of-the-art results, with the novel solver based on approximate mean-point correspondences being more robust and accurate than the affine-based solver. Charalambos Tzamos, Viktor Kocur, Yaqing Ding 0001, Daniel Barath, Zuzana Berger Haladová, Torsten Sattler, Zuzana Kukelova |
CVPR | 4 |
| 2025 | DepthSplat: Connecting Gaussian Splatting and DepthabstractGaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present Depth-Splat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pretrained monocular depth features, leading to high-quality feed-forward 3D Gaussian splatting reconstructions. We also show that Gaussian splatting can serve as an unsupervised pre-training objective for learning powerful depth models from large-scale multi-view posed datasets. We validate the synergy between Gaussian splatting and depth estimation through extensive ablation and cross-task transfer experiments. Our DepthSplat achieves state-of-the-art performance on ScanNet, RealEstate10K and DL3DV datasets in terms of both depth estimation and novel view synthesis, demonstrating the mutual benefits of connecting both tasks. In addition, DepthSplat enables feed-forward reconstruction from 12 input views (512 × 960 resolutions) in 0.6 seconds. Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger 0001, Marc Pollefeys |
CVPR | 5 |
| 2025 | HouseTour: A Virtual Real Estate A(I)gentabstractWe introduce HouseTour, a method for spatially-aware 3D camera trajectory and natural language summary generation from a collection of images depicting an existing 3D space. Unlike existing vision-language models (VLMs), which struggle with geometric reasoning, our approach generates smooth video trajectories via a diffusion process constrained by known camera poses and integrates this information into the VLM for 3D-grounded descriptions. We synthesize the final video using 3D Gaussian splatting to render novel views along the trajectory. To support this task, we present the HouseTour dataset, which includes over 1,200 house-tour videos with camera poses, 3D reconstructions, and real estate descriptions. Experiments demonstrate that incorporating 3D camera trajectories into the text generation process improves performance over methods handling each task independently. We evaluate both individual and end-to-end performance, introducing a new joint metric. Our work enables automated, professional-quality video creation for real estate and touristic applications without requiring specialized expertise or equipment. Ata Çelen, Marc Pollefeys, Daniel Barath, Iro Armeni |
ICCV | 3 |
| 2025 | Planar Affine Rectification from Local Change of Scale and Orientation
Yuval Nissan, Marc Pollefeys, Daniel Barath |
ICCV | 3 |
| 2025 | Hierarchical 3D Scene Graphs Construction Outdoors
Jon Nyffeler, Federico Tombari, Daniel Barath |
ICCV | 3 |
| 2025 | Object-X: Learning to Reconstruct Multi-Modal 3D Object RepresentationsabstractLearning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics.
Existing methods often rely on task-specific embeddings that are tailored either for semantic understanding or geometric reconstruction.
As a result, these embeddings typically cannot be decoded into explicit geometry and simultaneously reused across tasks.
In this paper, we propose Object-X, a versatile multi-modal object representation framework capable of encoding rich object embeddings (e.g., images, point cloud, text) and decoding them back into detailed geometric and visual reconstructions.
Object-X operates by geometrically grounding the captured modalities in a 3D voxel grid and learning an unstructured embedding fusing the information from the voxels with the object attributes.
The learned embedding enables 3D Gaussian Splatting-based object reconstruction, while also supporting a range of downstream tasks, including scene alignment, single-image 3D object reconstruction, and localization.
Evaluations on two challenging real-world datasets demonstrate that Object-X produces high-fidelity novel-view synthesis comparable to standard 3D Gaussian Splatting, while significantly improving geometric accuracy.
Moreover, Object-X achieves competitive performance with specialized methods in scene alignment and localization.
Critically, our object-centric descriptors require 3-4 orders of magnitude less storage compared to traditional image- or point cloud-based approaches, establishing Object-X as a scalable and highly practical solution for multi-modal 3D scene representation. Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath |
NeurIPS | 4 |
| 2025 | Breaking the Frame: Visual Place Recognition by Overlap PredictionabstractVisual place recognition methods struggle with occlusion and partial visual overlaps. We propose a novel visual place recognition approach based on overlap prediction, called VOP, shifting from traditional reliance on global image similarities and local features to image overlap prediction. VOP proceeds co-visible image sections by obtaining patch-level embeddings using a Vision Transformer backbone and establishing patch-to-patch correspondences without requiring expensive feature detection and matching. Our approach uses a voting mechanism to assess overlap scores for potential database images. It provides a nuanced image retrieval metric in challenging scenarios. Experimental results show that VOP leads to more accurate relative pose estimation and localization results on the retrieved image pairs than state-of-the-art base-lines on a number of large-scale, real-world indoor and outdoor benchmarks. The code is available at https://github.com/weitong8591/vop.git. Tong Wei 0002, Philipp Lindenberger, Jiri Matas, Daniel Barath |
WACV | 4 |
| 2025 | Multi-HexPlanes: A Lightweight Map Representation for Rendering and 3D ReconstructionabstractCreating maps of the world around us is paramount to many applications, including those related to robotics, such as navigation and inspection. Given the computational re-source limitations typical of robotic platforms, there is a pressing need for lightweight 3D representations that capture detailed texture and geometric information with min-imal storage. Traditional voxel-based approaches require substantial memory resources. On the other hand, neural implicit and 3D Gaussian splatting representations require sig-nificant computational power (GPUs) and can hardly run in real time. In this paper, we introduce a novel scene representation, Multi-HexPlanes, that divides 3D environments into large boxes and utilizes the faces of the boxes to encapsulate texture and geometric information. This representation re-duces the memory requirement to store the map, making our approach especially suitable for systems with limited memory. Through extensive evaluations on large-scale datasets, we find that our method achieves better performance on rendering and more complete 3D reconstruction. We also demonstrate that our map representation can output dense feature points with rich geometric information for down-stream tasks, such as training 3D Gaussian splats. The proposed technique promises substantial improvements in real-time 3D mapping applications, particularly for devices constrained by processing power and storage. Jianhao Zheng, Gábor Valasek, Daniel Barath, Iro Armeni |
WACV | 3 |
| 2025 | Generalized Relative Pose and Scale from Affine Correspondences
Wanting Xu, Marc Pollefeys, Daniel Barath, Laurent Kneip |
Int. J. Comput. Vis. | 4 |
| 2024 | Handbook on Leveraging Lines for Two-View Relative Pose EstimationabstractWe propose an approach for estimating the relative pose between calibrated image pairs by jointly exploiting points, lines, and their coincidences in a hybrid manner. We investigate all possible configurations where these data modalities can be used together and review the minimal solvers available in the literature. Our hybrid framework combines the advantages of all configurations, enabling robust and accurate estimation in challenging environments. In addition, we design a method for jointly estimating multiple vanishing point correspondences in two images, and a bundle adjustment that considers all relevant data modalities. Experiments on various indoor and outdoor datasets show that our approach outperforms point-based methods, improving AUC@ 10° by 1-7 points while running at comparable speeds. The source code of the solvers and hybrid framework will be made public. Petr Hruby, Shaohui Liu, Rémi Pautrat, Marc Pollefeys, Daniel Barath |
3DV | 5 |
| 2024 | Q-REG: End-to-End Trainable Point Cloud Registration with Surface CurvatureabstractPoint cloud registration has seen recent success with several learning-based methods that focus on correspondence matching, and as such, optimize only for this objective. Following the learning step of correspondence matching, they evaluate the estimated rigid transformation with a RANSAC-like framework. While it is an indispensable component of these methods, it prevents a fully end-to-end training, leaving the objective to minimize the pose error non-served. We present a novel solution, Q-REG, which utilizes rich geometric information to estimate the rigid pose from a single correspondence. Q-REG allows to formalize the robust estimation as an exhaustive search, hence enabling end-to-end training that optimizes over both objectives of correspondence matching and rigid pose estimation. We demonstrate in the experiments that Q-REG is agnostic to the correspondence matching method and provides consistent improvement both when used only in inference and in end-to-end training. It sets a new state-of-the-art on the 3DMatch, KITTI and ModelNet benchmarks. Our code and models are available at https://github.com/jinsz/Q-REG. Shengze Jin, Daniel Barath, Marc Pollefeys, Iro Armeni |
3DV | 2 |
| 2024 | Multiway Point Cloud Mosaicking with Diffusion and Global OptimizationabstractWe introduce a novel framework for multiway point cloud mosaicking (named Wednesday), designed to co-align sets of partially overlapping point clouds - typically obtained from 3D scanners or moving RGB-D cameras - into a unified coordinate system. At the core of our approach is ODIN, a learned pairwise registration algorithm that iteratively identifies overlaps and refines attention scores, employing a diffusion-based process for denoising pairwise correlation matrices to enhance matching accuracy. Further steps include constructing a pose graph from all point clouds, performing rotation averaging, a novel robust algorithm for re-estimating translations optimally in terms of consensus maximization and translation optimization. Finally, the point cloud rotations and positions are optimized jointly by a diffusion-based approach. Tested on four diverse, large-scale datasets, our method achieves state-of-the-art pairwise and multiway registration results by a large margin on all benchmarks. Our code and models are available at https://github.com/jinsz/Multiway-Point-Cloud-Mosaicking-with-Diffusion-and-Global-Optimization. Shengze Jin, Iro Armeni, Marc Pollefeys, Daniel Barath |
CVPR | 4 |
| 2024 | Absolute Pose from One or Two Scaled and Oriented FeaturesabstractKeypoints used for image matching often include an estimate of the feature scale and orientation. While recent work has demonstrated the advantages of using feature scales and orientations for relative pose estimation, relatively little work has considered their use for absolute pose estimation. We introduce minimal solutions for absolute pose from two oriented feature correspondences in the general case, or one scaled and oriented correspondence given a known vertical direction. Nowadays, assuming a known direction is not particularly restrictive as modern consumer devices, such as smartphones or drones, are equipped with Inertial Measurement Units (IMU) that provide the gravity direction by default. Compared to traditional absolute pose methods requiring three point correspondences, our solvers need a smaller minimal sample, reducing the cost and complexity of robust estimation. Evaluations on large-scale and public real datasets demonstrate the advantage of our methods for fast and accurate localization in challenging conditions. Code is available at https://github.com/danini/absolute-pose-from-oriented-and-sealed-features. Jonathan Ventura, Zuzana Kukelova, Torsten Sattler, Daniel Barath |
CVPR | 4 |
| 2024 | DGC-GNN: Leveraging Geometry and Color Cues for Visual Descriptor-Free 2D-3D MatchingabstractMatching 2D keypoints in an image to a sparse 3D point cloud of the scene without requiring visual descriptors has garnered increased interest due to its low memory requirements, inherent privacy preservation, and reduced need for expensive 3D model maintenance compared to visual descriptor-based methods. However, existing algorithms of-ten compromise on performance, resulting in a significant de-terioration compared to their descriptor-based counterparts. In this paper, we introduce DGC-GNN, a novel algorithm that employs a global-to-local Graph Neural Network (GNN) that progressively exploits geometric and color cues to rep-resent keypoints, thereby improving matching accuracy. Our procedure encodes both Euclidean and angular relations at a coarse level, forming the geometric embedding to guide the point matching. We evaluate DGC-GNN on both indoor and outdoor datasets, demonstrating that it not only doubles the accuracy of the state-of-the-art visual descriptor-free algorithm but also substantially narrows the performance gap between descriptor-based and descriptor-free methods.11The code and trained models are available at: https://github.com/AaltoVision/DGC-GNN-release. Shuzhe Wang, Juho Kannala, Daniel Barath |
CVPR | 3 |
| 2024 | StereoGlue: Robust Estimation with Single-Point Solvers
Daniel Barath, Dmytro Mishkin, Luca Cavalli, Paul-Edouard Sarlin, Petr Hruby, Marc Pollefeys |
ECCV (57) | 1 |
| 2024 | "Where am I?" Scene Retrieval with Language
Daniel Barath, Iro Armeni, Marc Pollefeys, Hermann Blum |
ECCV (37) | 2 |
| 2024 | Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization Using Geometrical Information
Luca Di Giammarino, Giorgio Grisetti, Marc Pollefeys, Hermann Blum, Daniel Barath |
ECCV (86) | 6 |
| 2024 | Semicalibrated Relative Pose from an Affine Correspondence and Monodepth
Petr Hruby, Marc Pollefeys, Daniel Barath |
ECCV (40) | 3 |
| 2024 | Learning to Make Keypoints Sub-pixel Accurate
Shinjeong Kim, Marc Pollefeys, Daniel Barath |
ECCV (78) | 3 |
| 2024 | SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs
Yang Miao 0003, Francis Engelmann, Olga Vysotska, Federico Tombari, Marc Pollefeys, Daniel Barath |
ECCV (8) | 6 |
| 2024 | Global Structure-from-Motion Revisited
Linfei Pan, Daniel Barath, Marc Pollefeys, Johannes L. Schönberger |
ECCV (40) | 2 |
| 2024 | Gravity-Aligned Rotation Averaging with Circular Regression
Linfei Pan, Marc Pollefeys, Daniel Barath |
ECCV (40) | 3 |
| 2024 | MAP-ADAPT: Real-Time Quality-Adaptive Semantic 3D Maps
Jianhao Zheng, Daniel Barath, Marc Pollefeys, Iro Armeni |
ECCV (39) | 2 |
| 2024 | Semantically Guided Feature Matching for Visual SLAMabstractWe introduce a new algorithm that utilizes semantic information to enhance feature matching in visual SLAM pipelines. The proposed method constructs a high-dimensional semantic descriptor for each detected ORB feature. When integrated with traditional visual ones, these descriptors aid in establishing accurate tentative point correspondences between consecutive frames. Additionally, our semantic descriptors enrich 3D map points, enhancing loop closure detection by providing deeper insights into the underlying map regions. Experiments on public large-scale datasets demonstrate that our technique surpasses the accuracy of established methods. Importantly, given its detector-agnostic nature, our algorithm also amplifies the efficacy of modern keypoint detectors, such as SuperPoint. The implementation of our algorithm can be found on Github3. Oguzhan Ilter, Iro Armeni, Marc Pollefeys, Daniel Barath |
ICRA | 4 |
| 2024 | Volumetric Semantically Consistent 3D Panoptic MappingabstractWe introduce an online 2D-to-3D semantic instance mapping algorithm aimed at generating comprehensive, accurate, and efficient semantic 3D maps suitable for autonomous agents in unstructured environments. The proposed approach is based on a Voxel-TSDF representation used in recent algorithms. It introduces novel ways of integrating semantic prediction confidence during mapping, producing semantic and instance-consistent 3D regions. Further improvements are achieved by graph optimization-based semantic labeling and instance refinement. The proposed method achieves accuracy superior to the state of the art on public large-scale datasets, improving on a number of widely used metrics. We also highlight a downfall in the evaluation of recent studies: using the ground truth trajectory as input instead of a SLAM-estimated one substantially affects the accuracy, creating a large gap between the reported results and the actual performance on real-world data. The code is available: https://github.com/y9miao/ConsistentPanopticSLAM. Yang Miao 0003, Iro Armeni, Marc Pollefeys, Daniel Barath |
IROS | 4 |
| 2023 | A Large-Scale Homography BenchmarkabstractWe present a large-scale dataset of Planes in 3D, Pi3D, of roughly 1000 planes observed in 10 000 images from the 1DSfM dataset, and HEB, a large-scale homography estimation benchmark leveraging Pi3D. The applications of the Pi3D dataset are diverse, e.g. training or evaluating monocular depth, surface normal estimation and image matching algorithms. The HEB dataset consists of 226 260 homographies and includes roughly 4M correspondences. The homographies link images that often undergo significant viewpoint and illumination changes. As applications of HEB, we perform a rigorous evaluation of a wide range of robust estimators and deep learning-based correspondence filtering methods, establishing the current state-of-the-art in robust homography estimation. We also evaluate the uncertainty of the SIFT orientations and scales w.r.t. the ground truth coming from the underlying homographies and provide codes for comparing uncertainty of custom detectors. The dataset is available at https://github.com/danini/homography-benchmark. Daniel Barath, Dmytro Mishkin, Michal Polic, Wolfgang Förstner, Jiri Matas |
CVPR | 1 |
| 2023 | Finding Geometric Models by Clustering in the Consensus SpaceabstractWe propose a new algorithm for finding an unknown number of geometric models, e.g., homographies. The problem is formalized as finding dominant model instances progressively without forming crisp point-to-model assignments. Dominant instances are found via a RANSAC-like sampling and a consolidation process driven by a model quality function considering previously proposed instances. New ones are found by clustering in the consensus space. This new formulation leads to a simple iterative algorithm with state-of-the-art accuracy while running in real-time on a number of vision problems - at least two orders of magnitude faster than the competitors on two-view motion estimation. Also, we propose a deterministic sampler reflecting the fact that real-world data tend to form spatially coherent structures. The sampler returns connected components in a progressively densified neighborhood-graph. We present a number of applications where the use of multiple geometric models improves accuracy. These include pose estimation from multiple generalized homographies; trajectory estimation of fast-moving objects; and we also propose a way of using multiple homographies in global SfM algorithms. Source code: https://github.com/danini/clustering-in-consensus-space. Daniel Barath, Denys Rozumnyi, Ivan Eichhardt, Levente Hajder, Jiri Matas |
CVPR | 1 |
| 2023 | DeepLSD: Line Segment Detection and Refinement with Deep Image GradientsabstractLine segments are ubiquitous in our human-made world and are increasingly used in vision tasks. They are complementary to feature points thanks to their spatial extent and the structural information they provide. Traditional line detectors based on the image gradient are extremely fast and accurate, but lack robustness in noisy images and challenging conditions. Their learned counterparts are more repeatable and can handle challenging images, but at the cost of a lower accuracy and a bias towards wireframe lines. We propose to combine traditional and learned approaches to get the best of both worlds: an accurate and robust line detector that can be trained in the wild without ground truth lines. Our new line segment detector, DeepLSD, processes images with a deep network to generate a line attraction field, before converting it to a surrogate image gradient magnitude and angle, which is then fed to any existing handcrafted line detector. Additionally, we propose a new optimization tool to refine line segments based on the attraction field and vanishing points. This refinement improves the accuracy of current deep detectors by a large margin. We demonstrate the performance of our method on low-level line detection metrics, as well as on several downstream tasks using multiple challenging datasets. The source code and models are available at https://github.com/cvg/DeepLSD. Rémi Pautrat, Daniel Barath, Viktor Larsson, Martin R. Oswald, Marc Pollefeys |
CVPR | 2 |
| 2023 | Revisiting Rotation Averaging: Uncertainties and Robust LossesabstractIn this paper, we revisit the rotation averaging problem applied in global Structure-from-Motion pipelines. We argue that the main problem of current methods is the minimized cost function that is only weakly connected with the input data via the estimated epipolar geometries. We propose to better model the underlying noise distributions by directly propagating the uncertainty from the point correspondences into the rotation averaging. Such uncertainties are obtained for free by considering the Jacobians of two-view refinements. Moreover, we explore integrating a variant of the MAGSAC loss into the rotation averaging problem, instead of using classical robust losses employed in current frameworks. The proposed method leads to results superior to baselines, in terms of accuracy, on large-scale public benchmarks. The code is public. https://github.com/zhangganlin/GlobalSfMpy Ganlin Zhang 0001, Viktor Larsson, Daniel Barath |
CVPR | 3 |
| 2023 | Fast Globally Optimal Surface Normal from an Affine CorrespondenceabstractWe present a new solver for estimating a surface normal from a single affine correspondence in two calibrated views. The proposed approach provides a new globally optimal solution for this over-determined problem and proves that it reduces to a linear system that can be solved extremely efficiently. This allows for performing significantly faster than other recent methods, solving the same problem and obtaining the same globally optimal solution. We demonstrate on 15k image pairs from standard benchmarks that the proposed approach leads to the same results as other optimal algorithms while being, on average, five times faster than the fastest alternative. Besides its theoretical value, we demonstrate that such an approach has clear benefits, e.g., in image-based visual localization, due to not requiring a dense point cloud to recover the surface normal. We show on the Cambridge Landmarks dataset that leveraging the proposed surface normal estimation further improves localization accuracy. Matlab and C++ implementations are also published in the supplementary material. Levente Hajder, Lajos Lóczi, Daniel Barath |
ICCV | 3 |
| 2023 | Vanishing Point Estimation in Uncalibrated Images with Prior Gravity DirectionabstractWe tackle the problem of estimating a Manhattan frame, i.e. three orthogonal vanishing points, and the unknown focal length of the camera, leveraging a prior vertical direction. The direction can come from an Inertial Measurement Unit that is a standard component of recent consumer devices, e.g., smartphones. We provide an exhaustive analysis of minimal line configurations and derive two new 2-line solvers, one of which does not suffer from singularities affecting existing solvers. Additionally, we design a new non-minimal method, running on an arbitrary number of lines, to boost the performance in local optimization. Combining all solvers in a hybrid robust estimator, our method achieves increased accuracy even with a rough prior. Experiments on synthetic and real-world datasets demonstrate the superior accuracy of our method compared to the state of the art, while having comparable runtimes. We further demonstrate the applicability of our solvers for relative rotation estimation. The code is available at https://github.com/cvg/VP-Estimation-with-Prior-Gravity. Rémi Pautrat, Shaohui Liu, Petr Hruby, Marc Pollefeys, Daniel Barath |
ICCV | 5 |
| 2023 | SGAligner: 3D Scene Alignment with Scene GraphsabstractBuilding 3D scene graphs has recently emerged as a topic in scene representation for several embodied AI applications to represent the world in a structured and rich manner. With their increased use in solving downstream tasks (e.g., navigation and room rearrangement), can we leverage and recycle them for creating 3D maps of environments, a pivotal step in agent operation? We focus on the fundamental problem of aligning pairs of 3D scene graphs whose overlap can range from zero to partial and can contain arbitrary changes. We propose SGAligner, the first method for aligning pairs of 3D scene graphs that is robust to in-the-wild scenarios (i.e., unknown overlap – if any – and changes in the environment). We get inspired by multimodality knowledge graphs and use contrastive learning to learn a joint, multi-modal embedding space. We evaluate on the 3RScan dataset and further showcase that our method can be used for estimating the transformation between pairs of 3D scenes. Since benchmarks for these tasks are missing, we create them on this dataset. The code, benchmark, and trained models are available on the project website. Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni |
ICCV | 4 |
| 2023 | P1AC: Revisiting Absolute Pose From a Single Affine CorrespondenceabstractAffine correspondences have traditionally been used to improve feature matching over wide baselines. While recent work has successfully used affine correspondences to solve various relative camera pose estimation problems, less attention has been given to their use in absolute pose estimation. We introduce the first general solution to the problem of estimating the pose of a calibrated camera given a single observation of an oriented point and an affine correspondence. The advantage of our approach (P1AC) is that it requires only a single correspondence, in comparison to the traditional point-based approach (P3P), significantly reducing the combinatorics in robust estimation. P1AC provides a general solution that removes restrictive assumptions made in prior work and is applicable to large-scale image-based localization. We propose a minimal solution to the P1AC problem and evaluate our novel solver on synthetic data, showing its numerical stability and performance under various types of noise. On standard image-based localization benchmarks we show that P1AC achieves more accurate results than the widely used P3P algorithm. Code for our method is available at https://github.com/jonathanventura/P1AC/. Jonathan Ventura, Zuzana Kukelova, Torsten Sattler, Daniel Barath |
ICCV | 4 |
| 2023 | Guiding Local Feature Matching with Surface CurvatureabstractWe propose a new method, called curvature similarity extractor (CSE), for improving local feature matching across images. CSE calculates the curvature of the local 3D surface patch for each detected feature point in a viewpoint-invariant manner via fitting quadrics to predicted monocular depth maps. This curvature is then leveraged as an additional signal in feature matching with off-the-shelf matchers like SuperGlue and LoFTR. Additionally, CSE enables end-to-end joint training by connecting the matcher and depth predictor networks. Our experiments demonstrate on large-scale real-world datasets that CSE consistently improves the accuracy of state-of-the-art methods. Fine-tuning the depth prediction network further enhances the accuracy. The proposed approach achieves state-of-the-art results on the ScanNet dataset, showcasing the effectiveness of incorporating 3D geometric information into feature matching.1 Shuzhe Wang, Juho Kannala, Marc Pollefeys, Daniel Barath |
ICCV | 4 |
| 2023 | Adaptive Reordering Sampler with Neurally Guided MAGSACabstractWe propose a new sampler for robust estimators that always selects the sample with the highest probability of consisting only of inliers. After every unsuccessful iteration, the inlier probabilities are updated in a principled way via a Bayesian approach. The probabilities obtained by the deep network are used as prior (so-called neural guidance) inside the sampler. Moreover, we introduce a new loss that exploits, in a geometrically justifiable manner, the orientation and scale that can be estimated for any type of feature, e.g., SIFT or SuperPoint, to estimate two-view geometry. The new loss helps to learn higher-order information about the underlying scene geometry. Benefiting from the new sampler and the proposed loss, we combine the neural guidance with the state-of-the-art MAGSAC++. Adaptive Reordering Sampler with Neurally Guided MAGSAC (ARS-MAGSAC) is superior to the state-of-the-art in terms of accuracy and run-time on the PhotoTourism and KITTI datasets for essential and fundamental matrix estimation. The code and trained models are available at https://github.com/weitong8591/ars_magsac. Tong Wei 0002, Jiri Matas, Daniel Barath |
ICCV | 3 |
| 2023 | Generalized Differentiable RANSACabstractWe propose ▽-RANSAC, a generalized differentiable RANSAC that allows learning the entire randomized robust estimation pipeline. The proposed approach enables the use of relaxation techniques for estimating the gradients in the sampling distribution, which are then propagated through a differentiable solver. The trainable quality function marginalizes over the scores from all the models estimated within ▽-RANSAC to guide the network learning accurate and useful inlier probabilities or to train feature detection and matching networks. Our method directly maximizes the probability of drawing a good hypothesis, allowing us to learn better sampling distributions. We test ▽-RANSAC on various real-world scenarios on fundamental and essential matrix estimation, and 3D point cloud registration, outdoors and indoors, with handcrafted and learning-based features. It is superior to the state-of-the-art in terms of accuracy while running at a similar speed to its less accurate alternatives. The code and trained models are available at https://github.com/weitong8591/differentiable_ransac. Tong Wei 0002, Alexander Shekhovtsov 0001, Jiri Matas, Daniel Barath |
ICCV | 5 |
| 2023 | On Making SIFT Features Affine CovariantabstractAbstract An approach is proposed for recovering affine correspondences (ACs) from orientation- and scale-covariant, e.g., SIFT, features exploiting pre-estimated epipolar geometry. The method calculates the affine parameters consistent with the epipolar geometry from the point coordinates and the scales and rotations which the feature detector obtains. The proposed closed-form solver returns a single solution and is extremely fast, i.e., 0.5 $$\upmu $$ μ seconds on average. Possible applications include estimating the homography from a single upgraded correspondence and, also, estimating the surface normal for each correspondence found in a pre-calibrated image pair (e.g., stereo rig). As the second contribution, we propose a minimal solver that estimates the relative pose of a vehicle-mounted camera from a single SIFT correspondence with the corresponding surface normal obtained from, e.g., upgraded ACs. The proposed algorithms are tested both on synthetic data and on a number of publicly available real-world datasets. Using the upgraded features and the proposed solvers leads to a significant speed-up in the homography, multi-homography and relative pose estimation problems with better or comparable accuracy to the state-of-the-art methods. Daniel Barath |
Int. J. Comput. Vis. | 1 |
| 2023 | Minimal Solvers for Relative Pose Estimation of Multi-Camera Systems using Affine Correspondences
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer |
Int. J. Comput. Vis. | 3 |
| 2022 | Pose-graph via Adaptive Image Re-ordering
Daniel Barath, Jana Noskova, Ivan Eichhardt, Jiri Matas |
BMVC | 1 |
| 2022 | Relative Pose from a Calibrated and an Uncalibrated Smartphone ImageabstractIn this paper, we propose a new minimal and a non-minimal solver for estimating the relative camera pose together with the unknown focal length of the second camera. This configuration has a number of practical benefits, e.g., when processing large-scale datasets. Moreover, it is resistant to the typical degenerate cases of the traditional six-point algorithm. The minimal solver requires four point correspondences and exploits the gravity direction that the built-in IMU of recent smart devices recover. We also propose a linear solver that enables estimating the pose from a larger-than-minimal sample extremely efficiently which then can be improved by, e.g., bundle adjustment. The methods are tested on 35654 image pairs from publicly available real-world and new datasets. When combined with a recent robust estimator, they lead to results superior to the traditional solvers in terms of rotation, translation and focal length accuracy, while being notably faster. Yaqing Ding 0001, Daniel Barath, Jian Yang 0003, Zuzana Kukelova |
CVPR | 2 |
| 2022 | Learning to Find Good Models in RANSACabstractWe propose the Model Quality Network, MQ-Net in short, for predicting the quality, e.g. the pose error of essential matrices, of models generated inside RANSAC. It replaces the traditionally used scoring techniques, e.g., inlier counting of RANSAC, truncated loss of MSAC, and the marginalization-based loss of MAGSAC++. Moreover, Minimal samples Filtering Network (MF-Net) is proposed for the early rejection of minimal samples that likely lead to degenerate models or to ones that are inconsistent with the scene geometry, e.g., due to the chirality constraint. We show on 54450 image pairs from public real-world datasets that the proposed MQ-Net leads to results superior to the state-of-the-art in terms of accuracy by a large margin. The proposed MF-Net accelerates the fundamental matrix estimation by five times and significantly reduces the essential matrix estimation time while slightly improving accuracy as well. Also, we show experimentally that consensus maximization, i.e. inlier counting, is not an inherently good measure of the model quality for relative pose estimation. The code is at https://github.com/danini/learning-goad-models-in-ransac. Daniel Barath, Luca Cavalli, Marc Pollefeys |
CVPR | 1 |
| 2022 | Relative Pose from SIFT Features
Daniel Barath, Zuzana Kukelova |
ECCV (32) | 1 |
| 2022 | Space-Partitioning RANSAC
Daniel Barath, Gábor Valasek |
ECCV (32) | 1 |
| 2022 | NeFSAC: Neurally Filtered Minimal Samples
Luca Cavalli, Marc Pollefeys, Daniel Barath |
ECCV (32) | 3 |
| 2022 | Relative Pose Solvers using Monocular DepthabstractWe describe a novel approach for using deep-learned priors to estimate the pose of a camera and show that these priors can be efficiently and accurately used for robust relative pose estimation. We use an off-the-shelf monocular depth network to provide an estimation of up-to-scale depth per pixel, and propose three new methods for solving for relative pose as well as a new algorithm for homography estimation. The additional signal provided by the depths leads to efficient solvers that require fewer correspondences than traditional methods and provide accurate and robust pose estimation when combined with state-of-the-art robust estimators, e.g., Graph-Cut RANSAC. The algorithms are tested on more than 70,000 publicly available image pairs from the 1DSfM dataset. The accuracy of the proposed methods are comparable or better than the standard five-point algorithm, and the reduced number of necessary correspondences speed up the robust estimation procedure, sometimes by orders of magnitude. Daniel Barath, Chris Sweeney |
ICPR | 1 |
| 2022 | Graph-Cut RANSAC: Local Optimization on Spatially Coherent StructuresabstractWe propose Graph-Cut RANSAC, GC-RANSAC in short, a new robust geometric model estimation method where the local optimization step is formulated as energy minimization with binary labeling, applying the graph-cut algorithm to select inliers. The minimized energy reflects the assumption that geometric data often form spatially coherent structures - it includes both a unary component representing point-to-model residuals and a binary term promoting spatially coherent inlier-outlier labelling of neighboring points. The proposed local optimization step is conceptually simple, easy to implement, efficient with a globally optimal inlier selection given the model parameters. Graph-Cut RANSAC, equipped with "the bells and whistles" of USAC and MAGSAC++, was tested on a range of problems using a number of publicly available datasets for homography, 6D object pose, fundamental and essential matrix estimation. It is more geometrically accurate than state-of-the-art robust estimators, fails less often and runs faster or with speed similar to less accurate alternatives. The source code is available at https://github.com/danini/graph-cut-ransac. Daniel Barath, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Marginalizing Sample ConsensusabstractA new method for robust estimation, MAGSAC++, is proposed. It introduces a new model quality (scoring) function that does not make inlier-outlier decisions, and a novel marginalization procedure formulated as an M-estimation with a novel class of M-estimators (a robust kernel) solved by an iteratively re-weighted least squares procedure. Instead of the inlier-outlier threshold, it requires only its loose upper bound which can be chosen from a significantly wider range. Also, we propose a new termination criterion and a technique for selecting a set of inliers in a data-driven manner as a post-processing step after the robust estimation finishes. On a number of publicly available real-world datasets for homography, fundamental matrix fitting and relative pose, MAGSAC++ produces results superior to the state-of-the-art robust methods. It is more geometrically accurate, fails fewer times, and it is often faster. It is shown that MAGSAC++ is significantly less sensitive to the setting of the threshold upper bound than the other state-of-the-art algorithms to the inlier-outlier threshold. Therefore, it is easier to be applied to unseen problems and scenes without acquiring information by hand about the setting of the inlier-outlier threshold. The source code and examples both in C++ and Python are available at https://github.com/danini/magsac. Daniel Barath, Jana Noskova, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Image Stitching with Locally Shared Rotation AxisabstractWe consider the problem of stitching image sequences with cameras undergoing pure rotational motion. We leverage the assumption of a locally constant rotation axis, i.e., neighboring frames have a shared but unknown rotation axis. This assumption holds in many common image capturing scenarios, e.g., panoramic sweeping motions. Using this additional constraint, we develop techniques for three-view camera rotation estimation; a minimal solver for the two-view estimation with a known rotation axis; and a globally optimal robust estimator for the two-view case. We show on publicly available datasets that the proposed methods lead to camera rotation estimation superior to the state-of-the-art in terms of accuracy with comparable run-time. The source code will be made available. Daniel Barath, Yaqing Ding 0001, Zuzana Kukelova, Viktor Larsson |
3DV | 1 |
| 2021 | Efficient Initial Pose-Graph Generation for Global SfMabstractWe propose ways to speed up the initial pose-graph generation for global Structure-from-Motion algorithms. To avoid forming tentative point correspondences by FLANN and geometric verification by RANSAC, which are the most time-consuming steps of the pose-graph creation, we propose two new methods – built on the fact that image pairs usually are matched consecutively. Thus, candidate relative poses can be recovered from paths in the partly-built pose-graph. We propose a heuristic for the A*traversal, considering global similarity of images and the quality of the pose-graph edges. Given a relative pose from a path, descriptor-based feature matching is made "light-weight" by exploiting the known epipolar geometry. To speed up PROSAC-based sampling when RANSAC is applied, we propose a third method to order the correspondences by their inlier probabilities from previous estimations. The algorithms are tested on 402130 image pairs from the 1DSfM dataset and they speed up the feature matching 17 times and pose estimation 5 times. Source code: https://github.com/danini/pose-graph-initialization Daniel Barath, Dmytro Mishkin, Ivan Eichhardt, Ilia Shipachev, Jiri Matas |
CVPR | 1 |
| 2021 | Globally Optimal Relative Pose Estimation With Gravity PriorabstractSmartphones, tablets and camera systems used, e.g., in cars and UAVs, are typically equipped with IMUs (inertial measurement units) that can measure the gravity vector accurately. Using this additional information, the y-axes of the cameras can be aligned, reducing their relative orientation to a single degree-of-freedom. With this assumption, we propose a novel globally optimal solver, minimizing the algebraic error in the least squares sense, to estimate the relative pose in the over-determined case. Based on the epipolar constraint, we convert the optimization problem into solving two polynomials with only two unknowns. Also, a fast solver is proposed using the first-order approximation of the rotation. The proposed solvers are compared with the state-of-the-art ones on four real-world datasets with approx. 50000 image pairs in total. Moreover, we collected a dataset, by a smartphone, consisting of 10933 image pairs, gravity directions and ground truth 3D reconstructions. The source code and dataset are available at https://github.com/yaqding/opt_pose_gravity Yaqing Ding 0001, Daniel Barath, Jian Yang 0003, Hui Kong 0001, Zuzana Kukelova |
CVPR | 2 |
| 2021 | Calibrated and Partially Calibrated Semi-Generalized HomographiesabstractIn this paper, we propose the first minimal solutions for estimating the semi-generalized homography given a perspective and a generalized camera. The proposed solvers use five 2D-2D image point correspondences induced by a scene plane. One group of solvers assumes the perspective camera to be fully calibrated, while the other estimates the unknown focal length together with the absolute pose parameters. This setup is particularly important in structure-from-motion and visual localization pipelines, where a new camera is localized in each step with respect to a set of known cameras and 2D-3D correspondences might not be available. Thanks to a clever parametrization and the elimination ideal method, our solvers only need to solve a univariate polynomial of degree five or three, respectively a system of polynomial equations in two variables. All proposed solvers are stable and efficient as demonstrated by a number of synthetic and real-world experiments. Snehal Bhayani, Torsten Sattler, Daniel Barath, Patrik Beliansky, Janne Heikkilä, Zuzana Kukelova |
ICCV | 3 |
| 2021 | Minimal Solutions for Panoramic Stitching Given Gravity PriorabstractWhen capturing panoramas, people tend to align their cameras with the vertical axis, i.e., the direction of gravity. Moreover, modern devices, e.g. smartphones and tablets, are equipped with an IMU (Inertial Measurement Unit) that can measure the gravity vector accurately. Using this prior, the y-axes of the cameras can be aligned or assumed to be already aligned, reducing the relative orientation to 1-DOF (degree of freedom). Exploiting this assumption, we propose new minimal solutions to panoramic stitching of images taken by cameras with coinciding optical centers, i.e. undergoing pure rotation. We consider six practical camera configurations, from fully calibrated ones up to a camera with unknown fixed or varying focal length and with or without radial distortion. The solvers are tested both on synthetic scenes, on more than 500k real image pairs from the Sun360 dataset, and from scenes captured by us using two smartphones equipped with IMUs. The new solvers have similar or better accuracy than the state-of-the-art ones and outperform them in terms of processing time. Yaqing Ding 0001, Daniel Barath, Zuzana Kukelova |
ICCV | 2 |
| 2021 | Minimal Cases for Computing the Generalized Relative Pose using Affine CorrespondencesabstractWe propose three novel solvers for estimating the relative pose of a multi-camera system from affine correspondences (ACs). A new constraint is derived interpreting the relationship of ACs and the generalized camera model. Using the constraint, we demonstrate efficient solvers for two types of motions assumed. Considering that the cameras undergo planar motion, we propose a minimal solution using a single AC and a solver with two ACs to overcome the degenerate case. Also, we propose a minimal solution using two ACs with known vertical direction, e.g., from an IMU. Since the proposed methods require significantly fewer correspondences than state-of-the-art algorithms, they can be efficiently used within RANSAC for outlier removal and initial motion estimation. The solvers are tested both on synthetic data and on real-world scenes from the KITTI odometry benchmark. It is shown that the accuracy of the estimated poses is superior to the state-of-the-art techniques. Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer |
ICCV | 3 |
| 2021 | VSAC: Efficient and Accurate Estimator for H and FabstractWe present VSAC, a RANSAC-type robust estimator with a number of novelties. It benefits from the introduction of the concept of independent inliers that improves significantly the efficacy of the dominant plane handling and, also, allows near error-free rejection of incorrect models, without false positives. The local optimization process and its application is improved so that it is run on average only once. Further technical improvements include adaptive sequential hypothesis verification and efficient model estimation via Gaussian elimination. Experiments on four standard datasets show that VSAC is significantly faster than all its predecessors and runs on average in 1-2 ms, on a CPU. It is two orders of magnitude faster and yet as precise as MAGSAC++, the currently most accurate estimator of two-view geometry. In the repeated runs on EVD, HPatches, PhotoTourism, and Kusvod2 datasets, it never failed. Maksym Ivashechkin, Daniel Barath, Jiri Matas |
ICCV | 2 |
| 2021 | Pose Estimation for Vehicle-mounted Cameras via Horizontal and Vertical PlanesabstractWe propose novel solvers for estimating the egomotion of a calibrated camera mounted to a moving vehicle from a single affine correspondence via recovering special homographies. For the first, second and third classes of solvers, the sought plane is expected to be perpendicular to one of the camera axes. For the fourth class, the plane is orthogonal to the ground with unknown normal, e.g., it is a building facade. All methods are solved via a linear system with a small coefficient matrix, thus, being extremely efficient. Both the minimal and over-determined cases can be solved by the proposed solvers. They are tested on synthetic data and on publicly available real-world datasets. The novel methods are more accurate or comparable to the traditional algorithms and are faster when included in state-of-the-art robust estimators. The source code is publicly available[1]. Istan Gergo Gal, Daniel Barath, Levente Hajder |
ICRA | 2 |
| 2021 | Efficient Recovery of Multi-Camera Motion from Two Affine CorrespondencesabstractWe propose an efficient method to estimate the relative pose of a multi-camera system from a minimum of two affine correspondences (ACs). Our solution is novel as it computes the 6DOF relative pose by utilizing a first-order rotation approximation. We directly derive a single polynomial based on the constraint between ACs and the generalized camera model. Then a closed-form solution is found analytically and it produces an accurate relative pose estimation efficiently. Benefiting from the low number of exploited correspondences and the speed of the solver, it speeds up robust estimators, e.g. RANSAC, significantly. The proposed method is evaluated both on synthetic data and real-world image sequences from the KITTI benchmark. It is shown that the proposed solver is superior to the state-of-the-art algorithms in terms of accuracy. Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer |
ICRA | 3 |
| 2020 | Homography-Based Egomotion Estimation Using Gravity and SIFT Features
Yaqing Ding 0001, Daniel Barath, Zuzana Kukelova |
ACCV (1) | 2 |
| 2020 | MAGSAC++, a Fast, Reliable and Accurate Robust EstimatorabstractA new method for robust estimation, MAGSAC++1, is proposed. It introduces a new model quality (scoring) function that does not require the inlier-outlier decision, and a novel marginalization procedure formulated as an M-estimation with a novel class of M-estimators (a robust kernel) solved by an iteratively re-weighted least squares procedure. We also propose a new sampler, Progressive NAPSAC, for RANSAC-like robust estimators. Exploiting the fact that nearby points often originate from the same model in real-world data, it finds local structures earlier than global samplers. The progressive transition from local to global sampling does not suffer from the weaknesses of purely localized samplers. On six publicly available realworld datasets for homography and fundamental matrix fitting, MAGSAC++ produces results superior to the state-of-the-art robust methods. It is faster, more geometrically accurate and fails less often. Daniel Barath, Jana Noskova, Maksym Ivashechkin, Jiri Matas |
CVPR | 1 |
| 2020 | EPOS: Estimating 6D Pose of Objects With SymmetriesabstractWe present a new method for estimating the 6D pose of rigid objects with available 3D models from a single RGB input image. The method is applicable to a broad range of objects, including challenging ones with global or partial symmetries. An object is represented by compact surface fragments which allow handling symmetries in a systematic manner. Correspondences between densely sampled pixels and the fragments are predicted using an encoder-decoder network. At each pixel, the network predicts: (i) the probability of each object's presence, (ii) the probability of the fragments given the object's presence, and (iii) the precise 3D location on each fragment. A data-dependent number of corresponding 3D locations is selected per pixel, and poses of possibly multiple object instances are estimated using a robust and efficient variant of the PnP-RANSAC algorithm. In the BOP Challenge 2019, the method outperforms all RGB and most RGB-D and D methods on the T-LESS and LM-O datasets. On the YCB-V dataset, it is superior to all competitors, with a large margin over the second-best RGB method. Source code is at: cmp.felk.cvut.cz/epos. Tomas Hodan, Daniel Barath, Jiri Matas |
CVPR | 2 |
| 2020 | Making Affine Correspondences Work in Camera Geometry Computation
Daniel Barath, Michal Polic, Wolfgang Förstner, Torsten Sattler, Tomás Pajdla, Zuzana Kukelova |
ECCV (11) | 1 |
| 2020 | Relative Pose from Deep Learned Depth and a Single Affine Correspondence
Ivan Eichhardt, Daniel Barath |
ECCV (12) | 2 |
| 2020 | Least-squares Optimal Relative Planar Motion for Vehicle-mounted CamerasabstractA new closed-form solver is proposed minimizing the algebraic error optimally, in the least squares sense, to estimate the relative planar motion of two calibrated cameras. The main objective is to solve the over-determined case, i.e., when a larger-than-minimal sample of point correspondences is given - thus, estimating the motion from at least three correspondences. The algorithm requires the camera movement to be constrained to a plane, e.g. mounted to a vehicle, and the image plane to be orthogonal to the ground.1The solver obtains the motion parameters as the roots of a 6th degree polynomial. It is validated both in synthetic experiments and on publicly available real-world datasets that using the proposed solver leads to results superior to the state-of-the-art in terms of geometric accuracy with no noticeable deterioration in the processing time. Levente Hajder, Daniel Barath |
ICRA | 2 |
| 2020 | Relative planar motion for vehicle-mounted cameras from a single affine correspondenceabstractTwo solvers are proposed for estimating the extrinsic camera parameters from a single affine correspondence assuming general planar motion. In this case, the camera movement is constrained to a plane and the image plane is orthogonal to the ground. The algorithms do not assume other constraints, e.g. the non-holonomic one, to hold. A new minimal solver is proposed for the semi-calibrated case, i.e. the camera parameters are known except a common focal length. Another method is proposed for the fully calibrated case. Due to requiring a single correspondence, robust estimation, e.g. histogram voting, leads to a fast and accurate procedure. The proposed methods are tested in our synthetic environment and on publicly available real datasets consisting of videos through tens of kilometers. They are superior to the state-of-the-art both in terms of accuracy and processing time. Levente Hajder, Daniel Barath |
ICRA | 2 |
| 2019 | Optimal Multi-view Correction of Local Affine Frames
Ivan Eichhardt, Daniel Barath |
BMVC | 2 |
| 2019 | MAGSAC: Marginalizing Sample ConsensusabstractA method called, sigma-consensus, is proposed to eliminate the need for a user-defined inlier-outlier threshold in RANSAC. Instead of estimating the noise sigma, it is marginalized over a range of noise scales. The optimized model is obtained by weighted least-squares fitting where the weights come from the marginalization over sigma of the point likelihoods of being inliers. A new quality function is proposed not requiring sigma and, thus, a set of inliers to determine the model quality. Also, a new termination criterion for RANSAC is built on the proposed marginalization approach. Applying sigma-consensus, MAGSAC is proposed with no need for a user-defined sigma and improving the accuracy of robust estimation significantly. It is superior to the state-of-the-art in terms of geometric accuracy on publicly available real-world datasets for epipolar geometry (F and E) and homography estimation. In addition, applying sigma-consensus only once as a post-processing step to the RANSAC output always improved the model quality on a wide range of vision problems without noticeable deterioration in processing time, adding a few milliseconds. Daniel Barath, Jiri Matas, Jana Noskova |
CVPR | 1 |
| 2019 | Homography From Two Orientation- and Scale-Covariant FeaturesabstractThis paper proposes a geometric interpretation of the angles and scales which the orientation- and scale-covariant feature detectors, e.g. SIFT, provide. Two new general constraints are derived on the scales and rotations which can be used in any geometric model estimation tasks. Using these formulas, two new constraints on homography estimation are introduced. Exploiting the derived equations, a solver for estimating the homography from the minimal number of two correspondences is proposed. Also, it is shown how the normalization of the point correspondences affects the rotation and scale parameters, thus achieving numerically stable results. Due to requiring merely two feature pairs, robust estimators, e.g. RANSAC, do significantly fewer iterations than by using the four-point algorithm. When using covariant features, e.g. SIFT, the information about the scale and orientation is given at no cost. The proposed homography estimation method is tested in a synthetic environment and on publicly available real-world datasets. Daniel Barath, Zuzana Kukelova |
ICCV | 1 |
| 2019 | Progressive-X: Efficient, Anytime, Multi-Model Fitting AlgorithmabstractThe Progressive-X algorithm, Prog-X in short, is proposed for geometric multi-model fitting. The method interleaves sampling and consolidation of the current data interpretation via repetitive hypothesis proposal, fast rejection, and integration of the new hypothesis into the kept instance set by labeling energy minimization. Due to exploring the data progressively, the method has several beneficial properties compared with the state-of-the-art. First, a clear criterion, adopted from RANSAC, controls the termination and stops the algorithm when the probability of finding a new model with a reasonable number of inliers falls below a threshold. Second, Prog-X is an any-time algorithm. Thus, whenever is interrupted, e.g. due to a time limit, the returned instances cover real and, likely, the most dominant ones. The method is superior to the state-of-the-art in terms of accuracy in both synthetic experiments and on publicly available real-world datasets for homography, two-view motion, and motion segmentation. Daniel Barath, Jiri Matas |
ICCV | 1 |
| 2019 | Optimal Multi-View Surface Normal Estimation Using Affine CorrespondencesabstractAn optimal, in the least squares sense, method is proposed to estimate surface normals in both stereo and multi-view cases. The proposed algorithm exploits exclusively photometric information via affine correspondences and estimates the normal for each correspondence independently. The normal is obtained as a root of a quartic polynomial. Therefore, the processing time is negligible. Eliminating the outliers, we propose a robust extension of the algorithm that combines maximum likelihood estimation and iteratively re-weighted least squares. The method has been validated on both synthetic and publicly available real-world datasets. It is superior to the state of the art in terms of accuracy and processing time. Besides, we demonstrate two possible applications: 1) using our algorithm as the seed-point generation step of patch-based multi-view stereo method, the obtained reconstruction is more accurate, and the error of the 3D points is reduced by 30% on average and 2) multi-plane fitting becomes more accurate applied to the resulting oriented point cloud. Daniel Barath, Ivan Eichhardt, Levente Hajder |
IEEE Trans. Image Process. | 1 |
| 2018 | Recovering Affine Features from Orientation- and Scale-Invariant Ones
Daniel Barath |
ACCV (1) | 1 |
| 2018 | Five-Point Fundamental Matrix Estimation for Uncalibrated CamerasabstractWe aim at estimating the fundamental matrix in two views from five correspondences of rotation invariant features obtained by e.g. the SIFT detector. The proposed minimal solver1 first estimates a homography from three correspondences assuming that they are co-planar and exploiting their rotational components. Then the fundamental matrix is obtained from the homography and two additional point pairs in general position. The proposed approach, combined with robust estimators like Graph-Cut RANSAC, is superior to other state-of-the-art algorithms both in terms of accuracy and number of iterations required. This is validated on synthesized data and 561 real image pairs. Moreover, the tests show that requiring three points on a plane is not too restrictive in urban environment and locally optimized robust estimators lead to accurate estimates even if the points are not entirely co-planar. As a potential application, we show that using the proposed method makes two-view multi-motion estimation more accurate. Daniel Barath |
CVPR | 1 |
| 2018 | Graph-Cut RANSACabstractA novel method for robust estimation, called Graph-Cut RANSAC1, GC-RANSAC in short, is introduced. To separate inliers and outliers, it runs the graph-cut algorithm in the local optimization (LO) step which is applied when a so-far-the-best model is found. The proposed LO step is conceptually simple, easy to implement, globally optimal and efficient. GC-RANSAC is shown experimentally, both on synthesized tests and real image pairs, to be more geometrically accurate than state-of-the-art methods on a range of problems, e.g. line fitting, homography, affine transformation, fundamental and essential matrix estimation. It runs in real-time for many problems at a speed approximately equal to that of the less accurate alternatives (in milliseconds on standard CPU). Daniel Barath, Jiri Matas |
CVPR | 1 |
| 2018 | Multi-class Model Fitting by Energy Minimization and Mode-Seeking
Daniel Barath, Jiri Matas |
ECCV (16) | 1 |
| 2018 | Efficient energy-based topological outlier rejection
Daniel Barath |
Comput. Vis. Image Underst. | 1 |
| 2018 | Efficient Recovery of Essential Matrix From Two Affine CorrespondencesabstractWe propose a method to estimate the essential matrix using two affine correspondences for a pair of calibrated perspective cameras. Two novel, linear constraints are derived between the essential matrix and a local affine transformation. The proposed method is also applicable to the over-determined case. We extend the normalization technique of Hartley to local affinities and show how the intrinsic camera matrices modify them. Even though perspective cameras are assumed, the constraints can straightforwardly be generalized to arbitrary camera models since they describe the relationship between local affinities and epipolar lines (or curves). Benefiting from the low number of exploited points, it can be used in robust estimators, e.g. RANSAC, as an engine, thus leading to significantly less iterations than the traditional point-based methods. The algorithm is validated both on synthetic and publicly available data sets and compared with the state-of-the-art. Its applicability is demonstrated on two-view multi-motion fitting, i.e., finding multiple fundamental matrices simultaneously, and outlier rejection. Daniel Barath, Levente Hajder |
IEEE Trans. Image Process. | 1 |
| 2017 | A Minimal Solution for Two-View Focal-Length Estimation Using Two Affine CorrespondencesabstractA minimal solution using two affine correspondences is presented to estimate the common focal length and the fundamental matrix between two semi-calibrated cameras - known intrinsic parameters except a common focal length. To the best of our knowledge, this problem is unsolved. The proposed approach extends point correspondence-based techniques with linear constraints derived from local affine transformations. The obtained multivariate polynomial system is efficiently solved by the hidden-variable technique. Observing the geometry of local affinities, we introduce novel conditions eliminating invalid roots. To select the best one out of the remaining candidates, a root selection technique is proposed outperforming the recent ones especially in case of high-level noise. The proposed 2-point algorithm is validated on both synthetic data and 104 publicly available real image pairs. A Matlab implementation of the proposed solution is included in the paper. Daniel Barath, Tekla Toth, Levente Hajder |
CVPR | 1 |
| 2017 | A theory of point-wise homography estimation
Daniel Barath, Levente Hajder |
Pattern Recognit. Lett. | 1 |
| 2016 | Accurate Closed-form Estimation of Local Affine Transformations Consistent with the Epipolar Geometry
Daniel Barath, Jiri Matas, Levente Hajder |
BMVC | 1 |
| 2016 | Multi-H: Efficient recovery of tangent planes in stereo images
Daniel Barath, Jiri Matas, Levente Hajder |
BMVC | 1 |
| 2016 | Energy-based topological outlier filteringabstractAn intuitive approach is proposed for outlier recognition among 2D point correspondences. The main novelty of the proposed method is the exploitation of feature point topology provided by Delaunay triangulation. The solution obtained by minimizing an energy originated from neighboring correspondences in order to remove incorrectly paired points. Assuming local, approximately rigid structures, it is able cope with non-rigid scenes. However, if the epipolar geometry is estimable, the additional information is exploited as well. The proposed method - called Delaunay Filtering - is validated on the publicly available AdelaideRMF dataset and outperforms the state-of-the-art, robust model-regression techniques. It is presented that it can be applied to image pairs for which epipolar geometry-based solutions fail. Daniel Barath, Levente Hajder |
ICPR | 1 |