Friedrich Fraundorfer

dblp:60/2234 · DBLP profile ↗
← Back
85ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0002-5805-8892ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 20 since 2021Systems, architecture and hardware · 20 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021
YearPublicationVenuePosition
2026 A Streamlined Attention-Based Network for Descriptor Extraction
abstract
We introduce SANDesc a Streamlined Attention-based Network for Descriptor extraction that aims to improve on existing architectures for keypoint description. Our descriptor network learns to compute descriptors that improve matching without modifying the underlying keypoint detector. We employ a revised U-Net-like architecture enhanced with Convolutional Block Attention Modules and residual paths, enabling effective local representation while maintaining computational efficiency. We refer to the building blocks of our model as Residual U-Net Blocks with Attention. The model is trained using a modified triplet loss in combination with a curriculum learning-inspired hard negative mining strategy, which improves training stability. Extensive experiments on HPatches, MegaDepth-1500, and the Image Matching Challenge 2021 show that training SANDESC on top of existing keypoint detectors leads to improved results on multiple matching tasks compared to the original keypoint descriptors. At the same time, SANDesc has a model complexity of just 2.4 million parameters. As a further contribution, we introduce a new urban dataset featuring 4K images and pre-calibrated intrinsics, designed to evaluate feature extractors. On this benchmark, SANDESC achieves substantial performance gains over the existing descriptors while operating with limited computational resources.
Mattia D'Urso, Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer
3DV6
2026 Autonomous Driving on Forest Roads under Unreliable GNSS*
Hamid Didari, Hans-Peter Wipfler, Simon Regner, Jakob Oberpertinger, Friedrich Fraundorfer, Gerald Steinbauer-Wagner
IV5
2025 PyTorchGeoNodes: Enabling Differentiable Shape Programs for 3D Shape Reconstruction
abstract
We propose PyTorchGeoNodes, a differentiable module for reconstructing 3D objects and their parameters from images using interpretable shape programs. Unlike traditional CAD model retrieval, shape programs allow reasoning about semantic parameters, editing, and a low memory footprint. Despite their potential, shape programs for 3D scene understanding have been largely overlooked. Our key contribution is enabling gradient-based optimization by parsing shape programs, or more precisely procedural models designed in Blender, into efficient PyTorch code. While there are many possible applications of our PyTochGeoNodes, we show that a combination of PyTorchGeoNodes with genetic algorithm is a method of choice to optimize both discrete and continuous shape program parameters for 3D reconstruction and understanding of 3D object parameters. Our modular framework can be further integrated with other reconstruction algorithms, and we demonstrate one such integration to enable procedural Gaussian splatting. Our experiments on the ScanNet dataset show that our method achieves accurate reconstructions while enabling, until now, unseen level of 3D scene understanding.
Sinisa Stekovic, Arslan Artykov, Stefan Ainetter, Mattia D'Urso, Friedrich Fraundorfer
CVPR5
2025 ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
abstract
Multi-view Stereo (MVS) aims to estimate depth and reconstruct 3D point clouds from a series of overlapping images. Recent learning-based MVS frameworks overlook the geometric information embedded in features and correlations, leading to weak cost matching. In this paper, we propose ICG-MVSNet, which explicitly integrates intra-view and cross-view relationships for depth estimation. Specifically, we develop an intra-view feature fusion module that leverages the feature coordinate correlations within a single image to enhance robust cost matching. Additionally, we introduce a lightweight cross-view aggregation module that efficiently utilizes the contextual information from volume correlations to guide regularization. Our method is evaluated on the DTU dataset and Tanks and Temples benchmark, consistently achieving competitive performance against state-of-the-art works, while requiring lower computational resources.
Jun Zhang 0102, Rafael Weilharter, Yuchen Rao, Kuangyi Chen, Runze Yuan, Friedrich Fraundorfer
ICME8
2025 EVLoc: Event-Based Visual Localization in LiDAR Maps via Event-Depth Registration
abstract
Event cameras are bio-inspired sensors with some notable features, including high dynamic range and low latency, which makes them exceptionally suitable for perception in challenging scenarios such as high-speed motion and extreme lighting conditions. In this paper, we explore their potential for localization within pre-existing LiDAR maps, a critical task for applications that require precise navigation and mobile manipulation. Our framework follows a paradigm based on the refinement of an initial pose. Specifically, we first project LiDAR points into 2D space based on a rough initial pose to obtain depth maps, and then employ an optical flow estimation network to align events with LiDAR points in 2D space, followed by camera pose estimation using a PnP solver. To enhance geometric consistency between these two inherently different modalities, we develop a novel frame-based event representation that improves structural clarity. Additionally, given the varying degrees of bias observed in the ground truth poses, we design a module that predicts an auxiliary variable as a regularization term to mitigate the impact of this bias on network convergence. Experimental results on several public datasets demonstrate the effectiveness of our proposed method. To facilitate future research, both the code and the pre-trained models are made available online11https://github.com/EasonChen99/EVLoc.
Kuangyi Chen, Jun Zhang 0102, Friedrich Fraundorfer
ICRA3
2025 MCTS With Refinement for Proposals Selection Games in Scene Understanding
abstract
We propose a novel method applicable in many scene understanding problems that adapts the Monte Carlo Tree Search (MCTS) algorithm, originally designed to learn to play games of high-state complexity. From a generated pool of proposals, our method jointly selects and optimizes proposals that minimize the objective term. In our first application for floor plan reconstruction from point clouds, our method selects and refines the room proposals, modelled as 2D polygons, by optimizing on an objective function combining the fitness as predicted by a deep network and regularizing terms on the room shapes. We also introduce a novel differentiable method for rendering the polygonal shapes of these proposals. Our evaluations on the recent and challenging Structured3D and Floor-SP datasets show significant improvements over the state-of-the-art both in speed and quality of reconstructions, without imposing hard constraints nor assumptions on the floor plan configurations. In our second application, we extend our approach to reconstruct general 3D room layouts from a color image and obtain accurate room layouts. We also show that our differentiable renderer can easily be extended for rendering 3D planar polygons and polygon embeddings. Our method shows high performance on the Matterport3D-Layout dataset, without introducing hard constraints on room layout configurations.
Sinisa Stekovic, Mahdi Rad, Friedrich Fraundorfer, Vincent Lepetit
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 HOC-Search: Efficient CAD Model and Pose Retrieval From RGB-D Scans
abstract
We present an automated and efficient approach for retrieving high-quality CAD models of objects and their poses in a scene captured by a moving RGB-D camera. We first investigate various objective functions to measure similarity between a candidate CAD object model and the available data, and the best objective function appears to be a ”render-and-compare” method comparing depth and mask rendering. We thus introduce a fast-search method that approximates an exhaustive search based on this objective function for simultaneously retrieving the object category, a CAD model, and the pose of an object given an approximate 3D bounding box. This method involves a search tree that organizes the CAD models and object properties including object category and pose for fast retrieval and an algorithm inspired by Monte Carlo Tree Search, that efficiently searches this tree. We show that this method retrieves CAD models that fit the real objects very well, with a speed-up factor of 10$\times$ to 120$\times$ compared to exhaustive search. We used our method to automatically retrieve CAD models for objects in the ScanNet dataset; our annotations are available at https://github.com/stefanainetter/SCANnotateDataset.
Stefan Ainetter, Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit
3DV3
2024 GMM-IKRS: Gaussian Mixture Models for Interpretable Keypoint Refinement and Scoring
Emanuele Santellani, Martin Zach, Christian Sormann, Mattia Rossi, Andreas Kuhn 0002, Friedrich Fraundorfer
ECCV (77)6
2024 SAda-Net: A Self-supervised Adaptive Stereo Estimation CNN For Remote Sensing Image Data
Dominik Hirner, Friedrich Fraundorfer
ICPR (10)2
2024 Robots Saving Lives: A Literature Review About Search and Rescue (SAR) in Harsh Environments
abstract
In recent years, the rise in both natural and man-made disasters, along with armed conflicts and terrorist threats, has elevated the demand for Search and Rescue (SAR) missions worldwide. This paper underscores the critical necessity to enhance SAR capacity, safety, and capabilities, with a primary goal of reducing response times through the integration of robots into SAR operations. The examination of research on robotized SAR highlights deficiencies in both software and hardware, particularly focusing on perception systems for robotized SAR platforms.
Kailin Tong, Berin Dikic, Selim Solmaz, Friedrich Fraundorfer, Daniel Watzenig
IV5
2024 HAMMER: Learning Entropy Maps to Create Accurate 3D Models in Multi-View Stereo
abstract
While the majority of recent Multi-View Stereo Networks estimates a depth map per reference image, their performance is then only evaluated on the fused 3D model obtained from all images. This approach makes a lot of sense since ultimately the point cloud is the result we are mostly interested in. On the flip side, it often leads to a burdensome manual search for the right fusion parameters in order to score well on the public benchmarks. In this work, we tackle the aforementioned problem with HAMMER, a Hierarchical And Memory-efficient MVSNet with Entropy-filtered Reconstructions. We propose to learn a filtering mask based on entropy, which, in combination with a simple two-view geometric verification, is sufficient to generate high quality 3D models of any input scene. Distinct from existing works, a tedious manual parameter search for the fusion step is not required. Furthermore, we take several precautions to keep the memory requirements for our method very low in the training as well as in the inference phase. Our method only requires 6 GB of GPU memory during training, while 3.6 GB are enough to process 1920×1024 images during inference. Experiments show that HAMMER ranks amongst the top published methods on the DTU and Tanks and Temples benchmarks in the official metrics, especially when keeping the fusion parameters fixed.
Rafael Weilharter, Friedrich Fraundorfer
WACV2
2023 S-TREK: Sequential Translation and Rotation Equivariant Keypoints for local feature extraction
abstract
In this work we introduce S-TREK, a novel local feature extractor that combines a deep keypoint detector, which is both translation and rotation equivariant by design, with a lightweight deep descriptor extractor. We train the S-TREK keypoint detector within a framework inspired by reinforcement learning, where we leverage a sequential procedure to maximize a reward directly related to keypoint repeatability. Our descriptor network is trained following a "detect, then describe" approach, where the descriptor loss is evaluated only at those locations where keypoints have been selected by the already trained detector. Extensive experiments on multiple benchmarks confirm the effectiveness of our proposed method, with S-TREK often outperforming other state-of-the-art methods in terms of repeatability and quality of the recovered poses, especially when dealing with in-plane rotations.
Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer
ICCV5
2023 Re: PolyWorld - A Graph Neural Network for Polygonal Scene Parsing
abstract
While most state-of-the-art instance segmentation methods produce pixel-wise segmentation masks, numerous applications demand precise vector polygons of detected objects instead of rasterized output. This paper proposes Re:PolyWorld as a remastered and improved version of PolyWorld, a neural network that extracts object vertices from an image and connects them optimally to generate precise polygons. The objective of this work was to overcome weaknesses and shortcomings of the original model, as well as introducing an improved polygonal representation to obtain a general-purpose method for polygon extraction in images. The architecture has been redesigned to not only exploit vertex features, but to also make use of the visual appearance of edges. To this end, an edge-aware Graph Neural Network predicts the connection strength between each pair of vertices, which is further used to compute the assignment by solving a differentiable optimal transport problem. The proposed redefinition of the polygonal scene turns the method into a powerful generalized approach that can be applied to a large variety of tasks and problem settings, such as building extraction, floorplan reconstruction and even wireframe parsing. Re:PolyWorld not only outperforms the original model on building extraction in aerial images, thanks to the proposed joint analysis of vertices and edges, but also beats the state-of-the-art in multiple other domains.
Stefano Zorzi, Friedrich Fraundorfer
ICCV2
2023 Automatically Annotating Indoor Images with CAD Models via RGB-D Scans
abstract
We present an automatic method for annotating images of indoor scenes with the CAD models of the objects by relying on RGB-D scans. Through a visual evaluation by 3D experts, we show that our method retrieves annotations that are at least as accurate as manual annotations, and can thus be used as ground truth without the burden of manually annotating 3D data. We do this using an analysis-by-synthesis approach, which compares renderings of the CAD models with the captured scene. We introduce a ’cloning procedure’ that identifies objects that have the same geometry, to annotate these objects with the same CAD models. This allows us to obtain complete annotations for the Scan-Net dataset and the recent ARKitScenes dataset. We will release these annotations publicly, as we believe they will be very useful for the computer vision community.
Stefan Ainetter, Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit
WACV3
2023 DELS-MVS: Deep Epipolar Line Search for Multi-View Stereo
abstract
We propose a novel approach for deep learning-based Multi-View Stereo (MVS). For each pixel in the reference image, our method leverages a deep architecture to search for the corresponding point in the source image directly along the corresponding epipolar line. We denote our method DELS-MVS: Deep Epipolar Line Search Multi-View Stereo. Previous works in deep MVS select a range of interest within the depth space, discretize it, and sample the epipolar line according to the resulting depth values: this can result in an uneven scanning of the epipolar line, hence of the image space. Instead, our method works directly on the epipolar line: this guarantees an even scanning of the image space and avoids both the need to select a depth range of interest, which is often not known a priori and can vary dramatically from scene to scene, and the need for a suitable discretization of the depth space. In fact, our search is iterative, which avoids the building of a cost volume, costly both to store and to process. Finally, our method performs a robust geometry-aware fusion of the estimated depth maps, leveraging a confidence predicted alongside each depth. We test DELS-MVS on the ETH3D, Tanks and Temples and DTU benchmarks and achieve competitive results with respect to state-of-the-art approaches.
Christian Sormann, Emanuele Santellani, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer
WACV5
2023 Minimal Solvers for Relative Pose Estimation of Multi-Camera Systems using Affine Correspondences
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
Int. J. Comput. Vis.4
2022 PolyWorld: Polygonal Building Extraction with Graph Neural Networks in Satellite Images
abstract
While most state-of-the-art instance segmentation methods produce binary segmentation masks, geographic and cartographic applications typically require precise vector polygons of extracted objects instead of rasterized output. This paper introduces PolyWorld, a neural network that directly extracts building vertices from an image and connects them correctly to create precise polygons. The model predicts the connection strength between each pair of vertices using a graph neural network and estimates the assignments by solving a differentiable optimal transport problem. Moreover, the vertex positions are optimized by minimizing a combined segmentation and polygonal angle difference loss. PolyWorld significantly outperforms the state of the art in building polygonization and achieves not only notable quantitative results, but also produces visually pleasing building polygons. Code and trained weights are publicly available at https://thub.com/zorzis/yWorl-PoldPretrainedNetwork.
Stefano Zorzi, Shabab Bazrafkan, Stefan Habenschuss, Friedrich Fraundorfer
CVPR4
2022 FCDSN-DC: An Accurate and Lightweight Convolutional Neural Network for Stereo Estimation with Depth Completion
abstract
We propose an accurate and lightweight convolutional neural network for stereo estimation with depth completion. We name this method fully-convolutional deformable similarity network with depth completion (FCDSN-DC). This method extends FC-DCNN by improving the feature extractor, adding a network structure for training highly accurate similarity functions and a network structure for filling inconsistent disparity estimates. The whole method consists of three parts. The first part consists of fully-convolutional densely connected layers that computes expressive features of rectified image pairs. The second part of our network learns highly accurate similarity functions between this learned features. It consists of densely-connected convolution layers with a deformable convolution block at the end to further improve the accuracy of the results. After this step an initial disparity map is created and the left-right consistency check is performed in order to remove inconsistent points. The last part of the network then uses this input together with the corresponding left RGB image in order to train a network that fills in the missing measurements. Consistent depth estimations are gathered around invalid points and are parsed together with the RGB points into a shallow CNN network structure in order to recover the missing values. We evaluate our method on challenging real world indoor and outdoor scenes, in particular Middlebury, KITTI and ETH3D were it produces competitive results. We furthermore show that this method generalizes well and is well suited for many applications without the need of further training. The code of our full framework is available at: https://github.com/thedodo/FCDSN-DC
Dominik Hirner, Friedrich Fraundorfer
ICPR2
2022 MD-Net: Multi-Detector for Local Feature Extraction
abstract
Establishing a sparse set of keypoint correspondences between images is a fundamental task in many computer vision pipelines. Often, this translates into a computationally expensive nearest neighbor search, where every keypoint descriptor at one image must be compared with all the descriptors at the others. In order to lower the computational cost of the matching phase, we propose a deep feature extraction network capable of detecting a predefined number of complementary sets of keypoints at each image. Since only the descriptors within the same set need to be compared across the different images, the matching phase computational complexity decreases with the number of sets. We train our network to predict the keypoints and compute the corresponding descriptors jointly. In particular, in order to learn complementary sets of keypoints, we introduce a novel unsupervised loss which penalizes intersections among the different sets. Additionally, we propose a novel descriptor-based weighting scheme meant to penalize the detection of keypoints with non-discriminative descriptors. With extensive experiments we show that our feature extraction network, trained only on synthetically warped images and in a fully unsupervised manner, achieves competitive results on 3D reconstruction and re-localization tasks at a reduced matching complexity.
Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer
ICPR5
2022 ATLAS-MVSNet: Attention Layers for Feature Extraction and Cost Volume Regularization in Multi-View Stereo
abstract
We present ATLAS-MVSNet, an end-to-end deep learning architecture relying on local attention layers for depth map inference from multi-view images. Distinct from existing works, we introduce a novel module design for neural networks, which we termed hybrid attention block, that utilizes the latest insights into attention in vision models. We are able to reap the benefits of attention in both, the carefully designed multi-stage feature extraction network and the cost volume regularization network. Our new approach displays significant improvement over its counterpart based purely on convolutions. While many state-of-the-art methods need multiple high-end GPUs in the training phase, we are able to train our network on a single consumer grade GPU. ATLAS-MVSNet exhibits excellent performance, especially in terms of accuracy, on the DTU dataset. Furthermore, ATLAS-MVSNet ranks amongst the top published methods on the online Tanks and Temples benchmark.
Rafael Weilharter, Friedrich Fraundorfer
ICPR2
2022 Relative Pose Estimation With a Single Affine Correspondence
abstract
In this article, we present four cases of minimal solutions for two-view relative pose estimation by exploiting the affine transformation between feature points, and we demonstrate efficient solvers for these cases. It is shown that under the planar motion assumption or with knowledge of a vertical direction, a single affine correspondence is sufficient to recover the relative camera pose. The four cases considered are two-view planar relative motion for calibrated cameras as a closed-form and least-squares solutions, a closed-form solution for unknown focal length, and the case of a known vertical direction. These algorithms can be used efficiently for outlier detection within a RANSAC loop and for initial motion estimation. All the methods are evaluated on both synthetic data and real-world datasets. The experimental results demonstrate that our methods outperform comparable state-of-the-art methods in accuracy with the benefit of a reduced number of needed RANSAC iterations. The source code is released at https://github.com/jizhaox/relative_pose_from_affine.
Banglei Guan, Ji Zhao 0001, Friedrich Fraundorfer
IEEE Trans. Cybern.5
2021 Depth-aware Object Segmentation and Grasp Detection for Robotic Picking Tasks
Stefan Ainetter, Christoph Böhm 0004, Rohit Dhakate, Stephan Weiss 0002, Friedrich Fraundorfer
BMVC5
2021 IB-MVS: An Iterative Algorithm for Deep Multi-View Stereo based on Binary Decisions
Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer
BMVC4
2021 Monte Carlo Scene Search for 3D Scene Understanding
abstract
We explore how a general AI algorithm can be used for 3D scene understanding to reduce the need for training data. More exactly, we propose a modification of the Monte Carlo Tree Search (MCTS) algorithm to retrieve objects and room layouts from noisy RGB-D scans. While MCTS was developed as a game-playing algorithm, we show it can also be used for complex perception problems. Our adapted MCTS algorithm has few easy-to-tune hyperparameters and can optimise general losses. We use it to optimise the posterior probability of objects and room layout hypotheses given the RGB-D data. This results in an analysis-by-synthesis approach that explores the solution space by rendering the current solution and comparing it to the RGB-D observations. To perform this exploration even more efficiently, we propose simple changes to the standard MCTS’ tree construction and exploration policy. We demonstrate our approach on the ScanNet dataset. Our method often retrieves configurations that are better than some manual annotations, especially on layouts.
Shreyas Hampali, Sinisa Stekovic, Sayan Deb Sarkar, Chetan Srinivasa Kumar, Friedrich Fraundorfer, Vincent Lepetit
CVPR5
2021 Minimal Cases for Computing the Generalized Relative Pose using Affine Correspondences
abstract
We propose three novel solvers for estimating the relative pose of a multi-camera system from affine correspondences (ACs). A new constraint is derived interpreting the relationship of ACs and the generalized camera model. Using the constraint, we demonstrate efficient solvers for two types of motions assumed. Considering that the cameras undergo planar motion, we propose a minimal solution using a single AC and a solver with two ACs to overcome the degenerate case. Also, we propose a minimal solution using two ACs with known vertical direction, e.g., from an IMU. Since the proposed methods require significantly fewer correspondences than state-of-the-art algorithms, they can be efficiently used within RANSAC for outlier removal and initial motion estimation. The solvers are tested both on synthetic data and on real-world scenes from the KITTI odometry benchmark. It is shown that the accuracy of the estimated poses is superior to the state-of-the-art techniques.
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
ICCV4
2021 MonteFloor: Extending MCTS for Reconstructing Accurate Large-Scale Floor Plans
abstract
We propose a novel method for reconstructing floor plans from noisy 3D point clouds. Our main contribution is a principled approach that relies on the Monte Carlo Tree Search (MCTS) algorithm to maximize a suitable objective function efficiently despite the complexity of the problem. Like previous work, we first project the input point cloud to a top view to create a density map and extract room proposals from it. Our method selects and optimizes the polygonal shapes of these room proposals jointly to fit the density map and outputs an accurate vectorized floor map even for large complex scenes. To do this, we adapt MCTS, an algorithm originally designed to learn to play games, to select the room proposals by maximizing an objective function combining the fitness with the density map as predicted by a deep network and regularizing terms on the room shapes. We also introduce a refinement step to MCTS that adjusts the shape of the room proposals. For this step, we propose a novel differentiable method for rendering the polygonal shapes of these proposals. We evaluate our method on the recent and challenging Structured3D and Floor-SP datasets and show a significant improvement over the state-of-the-art, without imposing any hard constraints nor assumptions on the floor plan configurations.
Sinisa Stekovic, Mahdi Rad, Friedrich Fraundorfer, Vincent Lepetit
ICCV3
2021 End-to-end Trainable Deep Neural Network for Robotic Grasp Detection and Semantic Segmentation from RGB
abstract
In this work, we introduce a novel, end-to-end trainable CNN-based architecture to deliver high quality results for grasp detection suitable for a parallel-plate gripper, and semantic segmentation. Utilizing this, we propose a novel refinement module that takes advantage of previously calculated grasp detection and semantic segmentation and further increases grasp detection accuracy. Our proposed network delivers state-of-the-art accuracy on two popular grasp dataset, namely Cornell and Jacquard. As additional contribution, we provide a novel dataset extension for the OCID dataset, making it possible to evaluate grasp detection in highly challenging scenes. Using this dataset, we show that semantic segmentation can additionally be used to assign grasp candidates to object classes, which can be used to pick specific objects in the scene.
Stefan Ainetter, Friedrich Fraundorfer
ICRA2
2021 Efficient Recovery of Multi-Camera Motion from Two Affine Correspondences
abstract
We propose an efficient method to estimate the relative pose of a multi-camera system from a minimum of two affine correspondences (ACs). Our solution is novel as it computes the 6DOF relative pose by utilizing a first-order rotation approximation. We directly derive a single polynomial based on the constraint between ACs and the generalized camera model. Then a closed-form solution is found analytically and it produces an accurate relative pose estimation efficiently. Benefiting from the low number of exploited correspondences and the speed of the solver, it speeds up robust estimators, e.g. RANSAC, significantly. The proposed method is evaluated both on synthetic data and real-world image sequences from the KITTI benchmark. It is shown that the proposed solver is superior to the state-of-the-art algorithms in terms of accuracy.
Banglei Guan, Ji Zhao 0001, Daniel Barath, Friedrich Fraundorfer
ICRA4
2021 End-to-End Semantic Segmentation and Boundary Regularization of Buildings from Satellite Imagery
abstract
Building footprint generation is a vital task of satellite imagery interpretation. However, the segmentation masks of buildings obtained by existing semantic segmentation networks often have blurred boundaries and irregular shapes. In this research, we propose a new boundary regularization network for building footprint generation in satellite images. More specifically, we consider semantic segmentation and boundary regularization in an end-to-end generative adversarial network (GAN). The learned building footprints are regularized by the interplay between the generator and discriminator. By doing so, the straight boundaries and geometric details of the building could be preserved. Experiments are conducted on a collected dataset of Planetscope satellite imagery (spatial resolution: 4.77 m/pixel). Our approach is much superior to the state-of-the-art methods in both quantitative and qualitative results.
Qingyu Li 0001, Stefano Zorzi, Yilei Shi, Friedrich Fraundorfer, Xiao Xiang Zhu 0001
IGARSS4
2020 DeepC-MVS: Deep Confidence Prediction for Multi-View Stereo Reconstruction
abstract
Deep Neural Networks (DNNs) have the potential to improve the quality of image-based 3D reconstructions. However, the use of DNNs in the context of 3D reconstruction from large and high-resolution image datasets is still an open challenge, due to memory and computational constraints. We propose a pipeline which takes advantage of DNNs to improve the quality of 3D reconstructions while being able to handle large and high-resolution datasets. In particular, we propose a confidence prediction network explicitly tailored for Multi-View Stereo (MVS) and we use it for both depth map outlier filtering and depth map refinement within our pipeline, in order to improve the quality of the final 3D reconstructions. We train our confidence prediction network on (semi-)dense ground truth depth maps from publicly available real world MVS datasets. With extensive experiments on popular benchmarks, we show that our overall pipeline can produce state-of-the-art 3D reconstructions, both qualitatively and quantitatively.
Andreas Kuhn 0005, Christian Sormann, Mattia Rossi, Oliver Erdler, Friedrich Fraundorfer
3DV5
2020 BP-MVSNet: Belief-Propagation-Layers for Multi-View-Stereo
abstract
In this work, we propose BP-MVSNet, a convolutional neural network (CNN)-based Multi-View-Stereo (MVS) method that uses a differentiable Conditional Random Field (CRF) layer for regularization. To this end, we propose to extend the BP layer [16] and add what is necessary to successfully use it in the MVS setting. We therefore show how we can calculate a normalization based on the expected 3D error, which we can then use to normalize the label jumps in the CRF. This is required to make the BP layer invariant to different scales in the MVS setting. In order to also enable fractional label jumps, we propose a differentiable interpolation step, which we embed into the computation of the pairwise term. These extensions allow us to integrate the BP layer into a multi-scale MVS network, where we continuously improve a rough initial estimate until we get high quality depth maps as a result. We evaluate the proposed BP-MVSNet in an ablation study and conduct extensive experiments on the DTU, Tanks and Temples and ETH3D data sets. The experiments show that we can significantly outperform the baseline and achieve state-of-the-art results.
Christian Sormann, Patrick Knöbelreiter, Andreas Kuhn 0005, Mattia Rossi, Thomas Pock, Friedrich Fraundorfer
3DV6
2020 Minimal Solutions for Relative Pose With a Single Affine Correspondence
abstract
In this paper we present four cases of minimal solutions for two-view relative pose estimation by exploiting the affine transformation between feature points and we demonstrate efficient solvers for these cases. It is shown, that under the planar motion assumption or with knowledge of a vertical direction, a single affine correspondence is sufficient to recover the relative camera pose. The four cases considered are two-view planar relative motion for calibrated cameras as a closed-form and a least-squares solution, a closed-form solution for unknown focal length and the case of a known vertical direction. These algorithms can be used efficiently for outlier detection within a RANSAC loop and for initial motion estimation. All the methods are evaluated on both synthetic data and real-world datasets from the KITTI benchmark. The experimental results demonstrate that our methods outperform comparable state-of-the-art methods in accuracy with the benefit of a reduced number of needed RANSAC iterations.
Banglei Guan, Ji Zhao 0001, Friedrich Fraundorfer
CVPR5
2020 Belief Propagation Reloaded: Learning BP-Layers for Labeling Problems
abstract
It has been proposed by many researchers that combining deep neural networks with graphical models can create more efficient and better regularized composite models. The main difficulties in implementing this in practice are associated with a discrepancy in suitable learning objectives as well as with the necessity of approximations for the inference. In this work we take one of the simplest inference methods, a truncated max-product Belief Propagation, and add what is necessary to make it a proper component of a deep learning model: connect it to learning formulations with losses on marginals and compute the backprop operation. This BP-Layer can be used as the final or an intermediate block in convolutional neural networks (CNNs), allowing us to design a hierarchical model composing BP inference and CNNs at different scale levels. The model is applicable to a range of dense prediction problems, is well-trainable and provides parameter-efficient and robust solutions in stereo, flow and semantic segmentation.
Patrick Knöbelreiter, Christian Sormann, Alexander Shekhovtsov 0001, Friedrich Fraundorfer, Thomas Pock
CVPR4
2020 General 3D Room Layout from a Single View by Render-and-Compare
Sinisa Stekovic, Shreyas Hampali, Mahdi Rad, Sayan Deb Sarkar, Friedrich Fraundorfer, Vincent Lepetit
ECCV (16)5
2020 Aerial Road Segmentation in the Presence of Topological Label Noise
abstract
The availability of large-scale annotated datasets has enabled Fully-Convolutional Neural Networks to reach outstanding performance on road extraction in aerial images. However, high-quality pixel-level annotation is expensive to produce and even manually labeled data often contains topological errors. Trading off quality for quantity, many datasets rely on already available yet noisy labels, for example from OpenStreetMap. In this paper, we explore the training of custom U-Nets built with ResNet and DenseNet backbones using noise-aware losses that are robust towards label omission and registration noise. We perform an extensive evaluation of standard and noise-aware losses, including a novel Bootstrapped DICE-Coefficient loss, on two challenging road segmentation benchmarks. Our losses yield a consistent improvement in overall extraction quality and exhibit a strong capacity to cope with severe label noise. Our method generalizes well to two other fine-grained topology delineation tasks: surface crack detection for quality inspection and cell membrane extraction in electron microscopy imagery.
Corentin Henry, Friedrich Fraundorfer, Eleonora Vig
ICPR2
2020 FC-DCNN: A densely connected neural network for stereo estimation
abstract
We propose a novel lightweight network for stereo estimation. Our network consists of a fully-convolutional densely connected neural network (FC-DCNN) that computes matching costs between rectified image pairs. Our FC-DCNN method learns expressive features and performs some simple but effective postprocessing steps. The densely connected layer structure connects the output of each layer to the input of each subsequent layer. This network structure and the fact that we do not use any fully-connected layers or 3D convolutions leads to a very lightweight network. The output of this network is used in order to calculate matching costs and create a cost-volume. Instead of using time and memory-inefficient cost-aggregation methods such as semiglobal matching or conditional random fields in order to improve the result, we rely on filtering techniques, namely median filter and guided filter. By computing a left-right consistency check we get rid of inconsistent values. Afterwards we use a watershed foreground-background segmentation on the disparity image with removed inconsistencies. This mask is then used to refine the final prediction. We show that our method works well for both challenging indoor and outdoor scenes by evaluating it on the Middlebury, KITTI and ETH3D benchmarks respectively. Our full framework is available at https://github.com/thedodo/FC-DCNN.
Dominik Hirner, Friedrich Fraundorfer
ICPR2
2020 Machine-learned Regularization and Polygonization of Building Segmentation Masks
abstract
We propose a machine learning based approach for automatic regularization and polygonization of building segmentation masks. Taking an image as input, we first predict building segmentation maps exploiting genericfullyconvolutionalnetwork(FCN). Agenerativeadversarialnetwork(GAN)is then involved to perform a regularization of building boundaries to make them more realistic, i.e., having more rectilinear outlines which construct right angles if required. This is achieved through the interplay between the discriminator which gives a probability of input image being true and generator that learns from discriminator's response to create more realistic images. Finally, we train the backboneconvolutionalneuralnetwork(CNN)which is adapted to predict sparse outcomes corresponding to building corners out of regularized building segmentation results. Experiments on three building segmentation datasets demonstrate that the proposed method is not only capable of obtaining accurate results, but also of producing visually pleasing building outlines parameterized as polygons.
Stefano Zorzi, Ksenia Bittner, Friedrich Fraundorfer
ICPR3
2020 Map-Repair: Deep Cadastre Maps Alignment and Temporal Inconsistencies Fix in Satellite Images
abstract
In the fast developing countries it is hard to trace new buildings construction or old structures destruction and, as a result, to keep the up-to-date cadastre maps. Moreover, due to the complexity of urban regions or inconsistency of data used for cadastre maps extraction, the errors in form of misalignment is a common problem. In this work, we propose an end-to-end deep learning approach which is able to solve inconsistencies between the input intensity image and the available building footprints by correcting label noises and, at the same time, misalignments if needed. The obtained results demonstrate the robustness of the proposed method to even severely misaligned examples that makes it potentially suitable for real applications, like OpenStreetMap correction.
Stefano Zorzi, Ksenia Bittner, Friedrich Fraundorfer
IGARSS3
2020 Casting Geometric Constraints in Semantic Segmentation as Semi-Supervised Learning
abstract
We propose a simple yet effective method to learn to segment new indoor scenes from video frames: State-of- the-art methods trained on one dataset, even as large as the SUNRGB-D dataset, can perform poorly when applied to images that are not part of the dataset, because of the dataset bias, a common phenomenon in computer vision. To make semantic segmentation more useful in practice, one can exploit geometric constraints. Our main contribution is to show that these constraints can be cast conveniently as semi-supervised terms, which enforce the fact that the same class should be predicted for the projections of the same 3D location in different images. This is interesting as we can exploit general existing techniques de- veloped for semi-supervised learning to efficiently incorporate the constraints. We show that this approach can efficiently and accurately learn to segment target sequences of ScanNet and our own target sequences using only annotations from SUNRGB-D, and geometric relations between the video frames of target sequences.
Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit
WACV2
2020 Comparison of monocular depth estimation methods using geometrically relevant metrics on the IBims-1 dataset
Tobias Koch 0002, Lukas Liebel, Marco Körner 0001, Friedrich Fraundorfer
Comput. Vis. Image Underst.4
2019 RESLAM: A real-time robust edge-based SLAM system
abstract
Simultaneous Localization and Mapping is a key requirement for many practical applications in robotics. In this work, we present RESLAM, a novel edge-based SLAM system for RGBD sensors. Due to their sparse representation, larger convergence basin and stability under illumination changes, edges are a promising alternative to feature-based or other direct approaches. We build a complete SLAM pipeline with camera pose estimation, sliding window optimization, loop closure and relocalisation capabilities that utilizes edges throughout all steps. In our system, we additionally refine the initial depth from the sensor, the camera poses and the camera intrinsics in a sliding window to increase accuracy. Further, we introduce an edge-based verification for loop closures that can also be applied for relocalisation. We evaluate RESLAM on wide variety of benchmark datasets that include difficult scenes and camera motions and also present qualitative results. We show that this novel edge-based SLAM system performs comparable to state-of-the-art methods, while running in real-time on a CPU. RESLAM is available as open-source software1.
Fabian Schenk, Friedrich Fraundorfer
ICRA2
2019 Regularization of Building Boundaries in Satellite Images Using Adversarial and Regularized Losses
abstract
In this paper we present a method for building boundary refinement and regularization in satellite images using a fully convolutional neural network trained with a combination of adversarial and regularized losses. Compared to a pure Mask R-CNN model, the overall algorithm can achieve equivalent performance in terms of accuracy and completeness. However, unlike Mask R-CNN that produces irregular footprints, our framework generates regularized and visually pleasing building boundaries which are beneficial in many applications.
Stefano Zorzi, Friedrich Fraundorfer
IGARSS2
2019 Force Field-Based Indirect Manipulation Of UAV Flight Trajectories
Werner Alexander Isop, Friedrich Fraundorfer
IROS2
2019 Rotational Alignment of IMU-camera Systems with 1-Point RANSAC
Banglei Guan, Ang Su, Friedrich Fraundorfer
PRCV (3)4
2019 Buildings Detection in VHR SAR Images Using Fully Convolution Neural Networks
abstract
This paper addresses the highly challenging problem of automatically detecting man-made structures especially buildings in very high-resolution (VHR) synthetic aperture radar (SAR) images. In this context, this paper has two major contributions. First, it presents a novel and generic workflow that initially classifies the spaceborne SAR tomography (TomoSAR) point clouds-generated by processing VHR SAR image stacks using advanced interferometric techniques known as TomoSAR-into buildings and nonbuildings with the aid of auxiliary information (i.e., either using openly available 2-D building footprints or adopting an optical image classification scheme) and later back project the extracted building points onto the SAR imaging coordinates to produce automatic large-scale benchmark labeled (buildings/nonbuildings) SAR data sets. Second, these labeled data sets (i.e., building masks) have been utilized to construct and train the state-of-the-art deep fully convolution neural networks with an additional conditional random field represented as a recurrent neural network to detect building regions in a single VHR SAR image. Such a cascaded formation has been successfully employed in computer vision and remote sensing fields for optical image classification but, to our knowledge, has not been applied to SAR images. The results of the building detection are illustrated and validated over a TerraSAR-X VHR spotlight SAR image covering approximately 39 km2-almost the whole city of Berlin- with the mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IEEE Trans. Geosci. Remote. Sens.3
2018 Semantically Aware Urban 3D Reconstruction with Plane-Based Regularization
Thomas Holzmann, Michael Maurer, Friedrich Fraundorfer, Horst Bischof
ECCV (14)3
2018 Visual Odometry Using a Homography Formulation with Decoupled Rotation and Translation Estimation Using Minimal Solutions
abstract
In this paper we present minimal solutions for two-view relative motion estimation based on a homography formulation. By assuming a known vertical direction (e.g. from an IMU) and assuming a dominant ground plane we demonstrate that rotation and translation estimation can be decoupled. This result allows us to reduce the number of point matches needed to compute a motion hypothesis. We then derive different algorithms based on this decoupling that allow an efficient estimation. We also demonstrate how these algorithms can be used efficiently to compute an optimal inlier set using exhaustive search or histogram voting instead of a traditional RANSAC step. Our methods are evaluated on synthetic data and on the KITTI data set, demonstrating that our methods are well suited for visual odometry in road driving scenarios.
Banglei Guan, Pascal Vasseur, Cédric Demonceaux, Friedrich Fraundorfer
ICRA4
2018 Extraction of Buildings in VHR SAR Images Using Fully Convolution Neural Networks
abstract
Modern spaceborne synthetic aperture radar (SAR) sensors, such as TerraSAR-X/TanDEM-X and COSMO-SkyMed, can deliver very high resolution (VHR) data beyond the inherent spatial scales (on the order of 1m) of buildings, constituting invaluable data source for large-scale urban mapping. Processing this VHR data with advanced interferometric techniques, such as SAR tomography (TomoSAR), enables the generation of 3-D (or even 4-D) TomoSAR point clouds from space. In this paper, we present a novel and generic workflow that exploits these TomoSAR point clouds in a way that is capable to automatically produce benchmark annotated (buildings/non-buildings) SAR datasets. These annotated datasets (building masks) have been utilized to construct and train the state-of-the-art deep Fully Convolution Neural Networks with an additional Conditional Random Field represented as a Recurrent Neural Network to detect building regions in a single VHR SAR image. The results of building detection are illustrated and validated over TerraSAR-X VHR spotlight SAR image covering approximately 39 km2- almost the whole city of Berlin - with mean pixel accuracies of around 93.84%.
Muhammad Shahzad 0002, Michael Maurer, Friedrich Fraundorfer, Yuanyuan Wang 0002, Xiao Xiang Zhu 0001
IGARSS3
2018 Building Detection and Segmentation Using a CNN with Automatically Generated Training Data
abstract
Significantly outperforming traditional machine learning methods, deep convolutional neural networks have gained increasing popularity in the application of image classification and segmentation. Nevertheless, deep learning-based methods usually require a large amount of training data, which is quite labor-intensive and time-demanding. To deal with the problem in generating training data, we propose in this paper a novel approach to generate image annotations by transferring labels from aerial images to UAV images and refine the annotations using a densely connected CRF model with an embedded naive Bayes classifier. The generated annotations not only present correct semantic labels, but also preserve accurate class boundaries. To validate the utility of these automatic annotations, we deploy them as training data for pixel-wise image segmentation and compare the results with the segmentation using manual annotations. Experiment results demonstrate that the automatic annotations can achieve comparable segmentation accuracy as the manual annotations.
Xiangyu Zhuo, Friedrich Fraundorfer, Franz Kurz, Peter Reinartz
IGARSS2
2018 Minimal solutions for the rotational alignment of IMU-camera systems using homography constraints
Banglei Guan, Friedrich Fraundorfer
Comput. Vis. Image Underst.3
2017 Primitive-based Surface Regularization for Urban 3D Reconstruction
Thomas Holzmann, Martin R. Oswald, Marc Pollefeys, Friedrich Fraundorfer, Horst Bischof
BMVC4
2017 Combining Edge Images and Depth Maps for Robust Visual Odometry
Fabian Schenk, Friedrich Fraundorfer
BMVC2
2017 Scalable Surface Reconstruction from Point Clouds with Extreme Scale and Density Diversity
abstract
In this paper we present a scalable approach for robustly computing a 3D surface mesh from multi-scale multi-view stereo point clouds that can handle extreme jumps of point density (in our experiments three orders of magnitude). The backbone of our approach is a combination of octree data partitioning, local Delaunay tetrahedralization and graph cut optimization. Graph cut optimization is used twice, once to extract surface hypotheses from local Delaunay tetrahedralizations and once to merge overlapping surface hypotheses even when the local tetrahedralizations do not share the same topology. This formulation allows us to obtain a constant memory consumption per sub-problem while at the same time retaining the density independent interpolation properties of the Delaunay-based optimization. On multiple public datasets, we demonstrate that our approach is highly competitive with the state-of-the-art in terms of accuracy, completeness and outlier resilience. Further, we demonstrate the multi-scale potential of our approach by processing a newly recorded dataset with 2 billion points and a point density variation of more than four orders of magnitude - requiring less than 9GB of RAM per process.
Christian Mostegel, Rudolf Prettenthaler, Friedrich Fraundorfer, Horst Bischof
CVPR3
2017 Robust edge-based visual odometry using machine-learned edges
abstract
In this work, we present a real-time robust edge-based visual odometry framework for RGBD sensors (REVO). Even though our method is independent of the edge detection algorithm, we show that the use of state-of-the-art machine-learned edges gives significant improvements in terms of robustness and accuracy compared to standard edge detection methods. In contrast to approaches that heavily rely on the photo-consistency assumption, edges are less influenced by lighting changes and the sparse edge representation offers a larger convergence basin while the pose estimates are also very fast to compute. Further, we introduce a measure for tracking quality, which we use to determine when to insert a new key frame. We show the feasibility of our system on real-world datasets and extensively evaluate on standard benchmark sequences to demonstrate the performance in a wide variety of scenes and camera motions. Our framework runs in real-time on the CPU of a laptop computer and is available online.
Fabian Schenk, Friedrich Fraundorfer
IROS2
2017 Evaluations on multi-scale camera networks for precise and geo-accurate reconstructions from aerial and terrestrial images with user guidance
Markus Rumpler, Alexander Tscharf, Christian Mostegel, Shreyansh Daftry, Christof Hoppe, Rudolf Prettenthaler, Friedrich Fraundorfer, Gerhard Mayer, Horst Bischof
Comput. Vis. Image Underst.7
2017 3D visual perception for self-driving cars using a multi-camera system: Calibration, mapping, localization, and obstacle detection
Christian Häne, Lionel Heng, Gim Hee Lee, Friedrich Fraundorfer, Paul Timothy Furgale, Torsten Sattler, Marc Pollefeys
Image Vis. Comput.4
2017 Homography Based Egomotion Estimation with a Common Direction
abstract
In this paper, we explore the different minimal solutions for egomotion estimation of a camera based on homography knowing the gravity vector between calibrated images. These solutions depend on the prior knowledge about the reference plane used by the homography. We then demonstrate that the number of matched points can vary from two to three and that a direct closed-form solution or a Gröbner basis based solution can be derived according to this plane. Many experimental results on synthetic and real sequences in indoor and outdoor environments show the efficiency and the robustness of our approach compared to standard methods.
Olivier Saurer, Pascal Vasseur, Rémi Boutteau, Cédric Demonceaux, Marc Pollefeys, Friedrich Fraundorfer
IEEE Trans. Pattern Anal. Mach. Intell.6
2016 Regularized 3D Modeling from Noisy Building Reconstructions
abstract
In this paper, we present a method for regularizing noisy 3D reconstructions, which is especially well suited for scenes containing planar structures like buildings. At horizontal structures, the input model is divided into slices and for each slice, an inside/outside labeling is computed. With the outlines of each slice labeling, we create an irregularly shaped volumetric cell decomposition of the whole scene. Then, an optimized inside/outside labeling of these cells is computed by solving an energy minimization problem. For the cell labeling optimization we introduce a novel smoothness term, where lines in the images are used to improve the regularization result. We show that our approach can take arbitrary dense meshed point clouds as input and delivers well regularized building models, which can be textured afterwards.
Thomas Holzmann, Friedrich Fraundorfer, Horst Bischof
3DV2
2016 Using Self-Contradiction to Learn Confidence Measures in Stereo Vision
abstract
Learned confidence measures gain increasing importance for outlier removal and quality improvement in stereo vision. However, acquiring the necessary training data is typically a tedious and time consuming task that involves manual interaction, active sensing devices and/or synthetic scenes. To overcome this problem, we propose a new, flexible, and scalable way for generating training data that only requires a set of stereo images as input. The key idea of our approach is to use different view points for reasoning about contradictions and consistencies between multiple depth maps generated with the same stereo algorithm. This enables us to generate a huge amount of training data in a fully automated manner. Among other experiments, we demonstrate the potential of our approach by boosting the performance of three learned confidence measures on the KITTI2012 dataset by simply training them on a vast amount of automatically generated training data rather than a limited amount of laser ground truth data.
Christian Mostegel, Markus Rumpler, Friedrich Fraundorfer, Horst Bischof
CVPR3
2016 Micro Aerial Projector - stabilizing projected images of an airborne robotics projection platform
abstract
A mobile flying projector is hard to build due to the limited size and payload capability of a micro aerial vehicle. Few flying projector designs have been studied in recent research. However, to date, no practical solution has been presented. We propose a versatile laser projection system enabling in-flight projection with feedforward correction for stabilization of projected images. We present a quantitative evaluation of the accuracy of the projection stabilization in two autonomous flight experiments. While this approach is our first step towards a flying projector, we foresee interesting applications, such as providing on-site instructions in various human machine interaction scenarios.
Werner Alexander Isop, Jesús Puerta Pestana, Gabriele Ermacora, Friedrich Fraundorfer, Dieter Schmalstieg
IROS4
2016 Aerial image sequence geolocalization with road traffic as invariant feature
Gellért Máttyus, Friedrich Fraundorfer
Image Vis. Comput.2
2014 A Homography Formulation to the 3pt Plus a Common Direction Relative Pose Problem
Olivier Saurer, Pascal Vasseur, Cédric Demonceaux, Friedrich Fraundorfer
ACCV (2)4
2014 Relative Pose Estimation for a Multi-camera System with Known Vertical Direction
abstract
In this paper, we present our minimal 4-point and linear 8-point algorithms to estimate the relative pose of a multi-camera system with known vertical directions, i.e. known absolute roll and pitch angles. We solve the minimal 4-point algorithm with the hidden variable resultant method and show that it leads to an 8-degree univariate polynomial that gives up to 8 real solutions. We identify a degenerated case from the linear 8-point algorithm when it is solved with the standard Singular Value Decomposition (SVD) method and adopt a simple alternative solution which is easy to implement. We show that our proposed algorithms can be efficiently used within RANSAC for robust estimation. We evaluate the accuracy of our proposed algorithms by comparisons with various existing algorithms for the multi-camera system on simulations and show the feasibility of our proposed algorithms with results from multiple real-world datasets.
Gim Hee Lee, Marc Pollefeys, Friedrich Fraundorfer
CVPR3
2013 Motion Estimation for Self-Driving Cars with a Generalized Camera
abstract
In this paper, we present a visual ego-motion estimation algorithm for a self-driving car equipped with a close-to-market multi-camera system. By modeling the multi-camera system as a generalized camera and applying the non-holonomic motion constraint of a car, we show that this leads to a novel 2-point minimal solution for the generalized essential matrix where the full relative motion including metric scale can be obtained. We provide the analytical solutions for the general case with at least one inter-camera correspondence and a special case with only intra-camera correspondences. We show that up to a maximum of 6 solutions exist for both cases. We identify the existence of degeneracy when the car undergoes straight motion in the special case with only intra-camera correspondences where the scale becomes unobservable and provide a practical alternative solution. Our formulation can be efficiently implemented within RANSAC for robust estimation. We verify the validity of our assumptions on the motion model by comparing our results on a large real-world dataset collected by a car equipped with 4 cameras with minimal overlapping field-of-views against the GPS/INS ground truth.
Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
CVPR2
2013 Robust pose-graph loop-closures with expectation-maximization
abstract
In this paper, we model the robust loop-closure pose-graph SLAM problem as a Bayesian network and show that it can be solved with the Classification Expectation-Maximization (EM) algorithm. In particular, we express our robust pose-graph SLAM as a Bayesian network where the robot poses and constraints are latent and observed variables. An additional set of latent variables is introduced as weights for the loop-constraints. We show that the weights can be chosen as the Cauchy function, which are iteratively computed from the errors between the predicted robot poses and observed loop-closure constraints in the Expectation step, and used to weigh the cost functions from the pose-graph loop-closure constraints in the Maximization step. As a result, outlier loop-closure constraints are assigned low weights and exert less influences in the pose-graph optimization within the EM iterations. To prevent the EM algorithm from getting stuck at local minima, we perform the EM algorithm multiple times where the loop constraints with very low weights are removed after each EM process. This is repeated until there are no more changes to the weights. We show proofs of the conceptual similarity between our EM algorithm and the M-Estimator. Specifically, we show that the weight function in our EM algorithm is equivalent to the robust residual function in the M-Estimator. We verify our proposed algorithm with experimental results from multiple simulated and real-world datasets, and comparisons with other existing works.
Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
IROS2
2013 Structureless pose-graph loop-closure with a multi-camera system on a self-driving car
abstract
In this paper, we propose a method to compute the pose-graph loop-closure constraints using multiple non/minimal overlapping field-of-views cameras mounted rigidly on a self-driving car without the need to reconstruct any 3D scene points. In particular, we show that the relative pose with metric scale between two loop-closing pose-graph vertices can be directly obtained from the epipolar geometry of the multicameras system. As a result, we avoid the additional time complexities and uncertainties from the reconstruction of 3D scene points which are needed by standard monocular and stereo approaches. In addition, there is a greater flexibility in choosing a configuration for the multi-camera system to cover a wider field-of-view so as to avoid missing out any loop-closure opportunities. We show that by expressing the point correspondences between two frames as Plücker lines and enforcing the planar motion constraint on the car, we are able to use multiple cameras as one and formulate the relative pose problem for loop-closure as a minimal problem which requires 3-point correspondences that yields up to six real solutions. The RANSAC algorithm is used to determine the correct solution and for robust estimation. We verify our method with results from multiple large-scale real-world data.
Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
IROS2
2013 Minimal Solutions for Pose Estimation of a Multi-Camera System
Gim Hee Lee, Bo Li 0018, Marc Pollefeys, Friedrich Fraundorfer
ISRR4
2013 Toward automated driving in cities using close-to-market sensors: An overview of the V-Charge Project
abstract
Future requirements for drastic reduction of CO2production and energy consumption will lead to significant changes in the way we see mobility in the years to come. However, the automotive industry has identified significant barriers to the adoption of electric vehicles, including reduced driving range and greatly increased refueling times. Automated cars have the potential to reduce the environmental impact of driving, and increase the safety of motor vehicle travel. The current state-of-the-art in vehicle automation requires a suite of expensive sensors. While the cost of these sensors is decreasing, integrating them into electric cars will increase the price and represent another barrier to adoption. The V-Charge Project, funded by the European Commission, seeks to address these problems simultaneously by developing an electric automated car, outfitted with close-to-market sensors, which is able to automate valet parking and recharging for integration into a future transportation system. The final goal is the demonstration of a fully operational system including automated navigation and parking. This paper presents an overview of the V-Charge system, from the platform setup to the mapping, perception, and planning sub-systems.
Paul Timothy Furgale, Ulrich Schwesinger, Martin Rufli, Wojciech Derendarz, Hugo Grimmett, Peter Mühlfellner, Stefan Wonneberger, Julian Timpner, Stephan Rottmann, Bo Li 0018, Bastian Schmidt, Thien-Nghia Nguyen, Elena Cardarelli, Stefano Cattani, Stefan Bruning, Sven Horstmann, Martin Stellmacher, Holger Mielenz, Kevin Köser, Markus Beermann, Christian Häne, Lionel Heng, Gim Hee Lee, Friedrich Fraundorfer, René Iser, Rudolph Triebel, Ingmar Posner, Paul Newman 0001, Lars C. Wolf, Marc Pollefeys, Stefan Brosig, Jan Effertz, Cédric Pradalier, Roland Siegwart
Intelligent Vehicles Symposium24
2012 SFly: Swarm of micro flying robots
abstract
The SFly project is an EU-funded project, with the goal to create a swarm of autonomous vision controlled micro aerial vehicles. The mission in mind is that a swarm of MAV's autonomously maps out an unknown environment, computes optimal surveillance positions and places the MAV's there and then locates radio beacons in this environment. The scope of the work includes contributions on multiple different levels ranging from theoretical foundations to hardware design and embedded programming. One of the contributions is the development of a new MAV, a hexacopter, equipped with enough processing power for onboard computer vision. A major contribution is the development of monocular visual SLAM that runs in real-time onboard of the MAV. The visual SLAM results are fused with IMU measurements and are used to stabilize and control the MAV. This enables autonomous flight of the MAV, without the need of a data link to a ground station. Within this scope novel analytical solutions for fusing IMU and vision measurements have been derived. In addition to the realtime local SLAM, an offline dense mapping process has been developed. For this the MAV's are equipped with a payload of a stereo camera system. The dense environment map is used to compute optimal surveillance positions for a swarm of MAV's. For this an optimiziation technique based on cognitive adaptive optimization has been developed. Finally, the MAV's have been equipped with radio transceivers and a method has been developed to locate radio beacons in the observed environment.
Markus Achtelik, Michael Achtelik, Yorick Brunet, Margarita Chli, Savvas A. Chatzichristofis, Jean-Dominique Decotignie, Klaus-Michael Doth, Friedrich Fraundorfer, Laurent Kneip, Daniel Gurdan, Lionel Heng, Elias B. Kosmatopoulos, Lefteris Doitsidis, Gim Hee Lee, Simon Lynen, Agostino Martinelli, Lorenz Meier, Marc Pollefeys, Damien Piguet, Alessandro Renzaglia, Davide Scaramuzza 0001, Roland Siegwart, Jan Stumpf, Petri Tanskanen, Chiara Troiani, Stephan Weiss 0002
IROS8
2012 Vision-based autonomous mapping and exploration using a quadrotor MAV
abstract
In this paper, we describe our autonomous vision-based quadrotor MAV system which maps and explores unknown environments. All algorithms necessary for autonomous mapping and exploration run on-board the MAV. Using a front-looking stereo camera as the main exteroceptive sensor, our quadrotor achieves these capabilities with both the Vector Field Histogram+ (VFH+) algorithm for local navigation, and the frontier-based exploration algorithm. In addition, we implement the Bug algorithm for autonomous wall-following which could optionally be selected as the substitute exploration algorithm in sparse environments where the frontier-based exploration under-performs. We incrementally build a 3D global occupancy map on-board the MAV. The map is used by the VFH+ and frontier-based exploration in dense environments, and the Bug algorithm for wall-following in sparse environments. During the exploration phase, images from the front-looking camera are transmitted over Wi-Fi to the ground station. These images are input to a large-scale visual SLAM process running off-board on the ground station. SLAM is carried out with pose-graph optimization and loop closure detection using a vocabulary tree. We improve the robustness of the pose estimation by fusing optical flow and visual odometry. Optical flow data is provided by a customized downward-looking camera integrated with a microcontroller while visual odometry measurements are derived from the front-looking stereo camera. We verify our approaches with experimental results.
Friedrich Fraundorfer, Lionel Heng, Dominik Honegger, Gim Hee Lee, Lorenz Meier, Petri Tanskanen, Marc Pollefeys
IROS1
2011 Autonomous obstacle avoidance and maneuvering on a vision-guided MAV using on-board processing
abstract
We present a novel stereo-based obstacle avoidance system on a vision-guided micro air vehicle (MAV) that is capable of fully autonomous maneuvers in unknown and dynamic environments. All algorithms run exclusively on the vehicle's on-board computer, and at high frequencies that allow the MAV to react quickly to obstacles appearing in its flight trajectory. Our MAV platform is a quadrotor aircraft equipped with an inertial measurement unit and two stereo rigs. An obstacle mapping algorithm processes stereo images, producing a 3D map representation of the environment; at the same time, a dynamic anytime path planner plans a collision-free path to a goal point.
Lionel Heng, Lorenz Meier, Petri Tanskanen, Friedrich Fraundorfer, Marc Pollefeys
ICRA4
2011 MAV visual SLAM with plane constraint
abstract
Bundle adjustment (BA) which produces highly accurate results for visual Simultaneous Localization and Mapping (SLAM) could not be used for Micro-Aerial Vehicles (MAVs) with limited processing power because of its O(N3) complexity. We observed that a consistent ground plane often exists for MAVs flying in both the indoor and outdoor urban environments. Therefore, in this paper, we propose a visual SLAM algorithm that make use of the plane constraint to reduce the complexity of BA. The reduction of complexity is achieved by refining only the current camera pose and most recent map points with BA that minimizes the reprojection errors and perpendicular distances between the most recent map points and the best fit plane with all the pre-existing map points. As a result, our algorithm is approximately constant time since the number of current camera pose and most recent map points remain approximately constant. In addition, the minimization of the perpendicular distances between the plane and map points would enforce consistency between the reconstructed map points and the actual ground plane.
Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
ICRA2
2011 PIXHAWK: A system for autonomous flight using onboard computer vision
abstract
We provide a novel hardware and software system for micro air vehicles (MAV) that allows high-speed, low-latency onboard image processing. It uses up to four cameras in parallel on a miniature rotary wing platform. The MAV navigates based on onboard processed computer vision in GPS-denied in- and outdoor environments. It can process in parallel images and inertial measurement information from multiple cameras for multiple purposes (localization, pattern recognition, obstacle avoidance) by distributing the images on a central, low-latency image hub. Furthermore the system can utilize low-bandwith radio links for communication and is designed and optimized to scale to swarm use. Experimental results show successful flight with a range of onboard computer vision algorithms, including localization, obstacle avoidance and pattern recognition.
Lorenz Meier, Petri Tanskanen, Friedrich Fraundorfer, Marc Pollefeys
ICRA3
2011 Real-time photo-realistic 3D mapping for micro aerial vehicles
abstract
In this paper, we proposed a method to recognize complex human daily activities including body activities and hand gestures simultaneously in an indoor environment. Three wearable motion sensors are attached to the right thigh, the waist, and the right hand of a person, while an optical motion capture system is used to obtain his/her location information. A three-level dynamic Bayesian network is implemented to model the intra-temporal and inter-temporal constraints among the location, body activity and hand gesture. The body activity and hand gesture are estimated using a Bayesian filter and the short-time Viterbi algorithm, which reduces the storage memory and the computational complexity. We conducted experiments in a mock apartment environment and the obtained results showed the effectiveness and accuracy of our algorithms.
Lionel Heng, Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
IROS3
2011 RS-SLAM: RANSAC sampling for visual FastSLAM
abstract
In this paper, we present our RS-SLAM algorithm for monocular camera where the proposal distribution is derived from the 5-point RANSAC algorithm and image feature measurement uncertainties instead of using the easily violated constant velocity model. We propose to do another RANSAC sampling within all the inliers that have the best RANSAC score to check for inlier misclassifications in the original correspondences and use all the hypotheses generated from these consensus sets in the proposal distribution. This is to mitigate data association errors (inlier misclassifications) caused by the observation that the consensus set from RANSAC that yields the highest score might not, in practice, contain all the true inliers due to noise on the feature measurements. Hypotheses which are less probable will eventually be eliminated in the particle filter resampling process. We also show in this paper that our monocular approach can be easily extended for stereo camera. Experimental results validate the potential of our approach.
Gim Hee Lee, Friedrich Fraundorfer, Marc Pollefeys
IROS2
2010 A Minimal Case Solution to the Calibrated Relative Pose Problem for the Case of Two Known Orientation Angles
Friedrich Fraundorfer, Petri Tanskanen, Marc Pollefeys
ECCV (4)1
2010 A benchmarking tool for MAV visual pose estimation
abstract
The large collections of datasets for researchers working on the Simultaneous Localization and Mapping problem are mostly collected from sensors such as wheel encoders and laser range finders mounted on ground robots. The recent growing interest in doing visual pose estimation with cameras mounted on micro-aerial vehicles however has made these datasets less useful. In this paper, we describe our work in creating new datasets collected from a sensor suite mounted on a quadrotor platform. Our sensor suite includes a forward looking camera, a downward looking camera, an inertial measurement unit and a Vicon system for groundtruth. We propose the use our datasets as benchmarking tools for future works on visual pose estimation for micro-aerial vehicles. We also show examples of how the datasets could be used for benchmarking visual pose estimation algorithms.
Gim Hee Lee, Markus Achtelik, Friedrich Fraundorfer, Marc Pollefeys, Roland Siegwart
ICARCV3
2010 Combining Monocular and Stereo Cues for Mobile Robot Localization Using Visual Words
abstract
This paper describes an approach for mobile robot localization using a visual word based place recognition approach. In our approach we exploit the benefits of a stereo camera system for place recognition. Visual words computed from SIFT features are combined with VIP (viewpoint invariant patches) features that use depth information from the stereo setup. The approach was evaluated under the ImageCLEF@ICPR 2010 competition. The results achieved on the competition datasets are published in this paper.
Friedrich Fraundorfer, Changchang Wu, Marc Pollefeys
ICPR1
2010 A constricted bundle adjustment parameterization for relative scale estimation in visual odometry
abstract
In this paper we address the problem of visual motion estimation (visual odometry) from a single vehicle mounted camera. One of the basic issues of visual odometry is relative scale estimation. We propose a method to compute the relative scales of a path by solving a bundle adjustment optimization problem. We introduce a constricted parameterization of the bundle adjustment problem, where only the distances between neighboring cameras are optimized, while the rotation angles and translation directions stay fixed. We will present visual odometry results for image data of a vehicle mounted onmidirectional camera for a track of 1000m length.
Friedrich Fraundorfer, Davide Scaramuzza 0001, Marc Pollefeys
ICRA1
2009 Absolute scale in structure from motion from a single vehicle mounted camera by exploiting nonholonomic constraints
abstract
In structure-from-motion with a single camera it is well known that the scene can be only recovered up to a scale. In order to compute the absolute scale, one needs to know the baseline of the camera motion or the dimension of at least one element in the scene. In this paper, we show that there exists a class of structure-from-motion problems where it is possible to compute the absolute scale completely automatically without using this knowledge, that is, when the camera is mounted on wheeled vehicles (e.g. cars, bikes, or mobile robots). The construction of these vehicles puts interesting constraints on the camera motion, which are known as “nonholonomic constraints”. The interesting case is when the camera has an offset to the vehicle's center of motion. We show that by just knowing this offset, the absolute scale can be computed with a good accuracy when the vehicle turns. We give a mathematical derivation and provide experimental results on both simulated and real data over a large image dataset collected during a 3 Km path. To our knowledge this is the first time nonholonomic constraints of wheeled vehicles are used to estimate the absolute scale. We believe that the proposed method can be useful in those research areas involving visual odometry and mapping with vehicle mounted cameras.
Davide Scaramuzza 0001, Friedrich Fraundorfer, Marc Pollefeys, Roland Siegwart
ICCV2
2009 Real-time monocular visual odometry for on-road vehicles with 1-point RANSAC
abstract
This paper presents a system capable of recovering the trajectory of a vehicle from the video input of a single camera at a very high frame-rate. The overall frame-rate is limited only by the feature extraction process, as the outlier removal and the motion estimation steps take less than 1 millisecond with a normal laptop computer. The algorithm relies on a novel way of removing the outliers of the feature matching process.We show that by exploiting the nonholonomic constraints of wheeled vehicles it is possible to use a restrictive motion model which allows us to parameterize the motion with only 1 feature correspondence. Using a single feature correspondence for motion estimation is the lowest model parameterization possible and results in the most efficient algorithms for removing outliers. Here we present two methods for outlier removal. One based on RANSAC and the other one based on histogram voting. We demonstrate the approach using an omnidirectional camera placed on a vehicle during a peak time tour in the city of Zurich. We show that the proposed algorithm is able to cope with the large amount of clutter of the city (other moving cars, buses, trams, pedestrians, sudden stops of the vehicle, etc.). Using the proposed approach, we cover one of the longest trajectories ever reported in real-time from a single omnidirectional camera and in cluttered urban scenes, up to 3 kilometers.
Davide Scaramuzza 0001, Friedrich Fraundorfer, Roland Siegwart
ICRA2
2009 Towards Large-Scale Visual Mapping and Localization
Marc Pollefeys, Jan-Michael Frahm, Friedrich Fraundorfer, Christopher Zach, Changchang Wu, Brian Clipp, David Gallup
ISRR3
2007 A Binning Scheme for Fast Hard Drive Based Image Search
abstract
In this paper we investigate how to scale a content based image retrieval approach beyond the RAM limits of a single computer and to make use of its hard drive to store the feature database. The feature vectors describing the images in the database are binned in multiple independent ways. Each bin contains images similar to a representative prototype. Each binning is considered through two stages of processing. First, the prototype closest to the query is found. Second, the bin corresponding to the closest prototype is fetched from disk and searched completely. The query process is repeatedly performing these two stages, each time with a binning independent of the previous ones. The scheme cuts down the hard drive access significantly and results in a major speed up. An experimental comparison between the binning scheme and a raw search shows competitive retrieval quality.
Friedrich Fraundorfer, Henrik Stewénius, David Nistér
CVPR1
2007 Topological mapping, localization and navigation using image collections
abstract
In this paper we present a highly scalable vision-based localization and mapping method using image collections. A topological world representation is created online during robot exploration by adding images to a database and maintaining a link graph. An efficient image matching scheme allows real-time mapping and global localization. The compact image representation allows us to create image collections containing millions of images, which enables mapping of very large environments. A path planning method using graph search is proposed and local geometric information is used to navigate in the topological map. Experiments show the good performance of the image matching for global localization and demonstrate path planning and navigation.
Friedrich Fraundorfer, Chris Engels, David Nistér
IROS1
2006 Piecewise planar scene reconstruction from sparse correspondences
Friedrich Fraundorfer, Konrad Schindler, Horst Bischof
Image Vis. Comput.1