Konrad Schindler

dblp:73/488 · DBLP profile ↗
← Back
139ranked-venue papers
11as first author
54since 2021 · last 2025
0000-0002-3172-9246ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 115 · 9 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 84 · 5 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
YearPublicationVenuePosition
2025 LoopSplat: Loop Closure by Registering 3D Gaussian Splats
abstract
Simultaneous Localization and Mapping (SLAM) based on 3D Gaussian Splats (3DGS) has recently shown promise towards more accurate, dense 3D scene maps. However, existing 3DGS-based methods fail to address the global consistency of the scene via loop closure and/or global bundle adjustment. To this end, we propose LoopSplat, which takes RGB-D images as input and performs dense mapping with 3DGS submaps and frame-to-model tracking. LoopSplat triggers loop closure online and computes relative loop edge constraints between submaps directly via 3DGS registration, leading to improvements in efficiency and accuracy over traditional global-to-local point cloud registration. It uses a robust pose graph optimization formulation and rigidly aligns the submaps to achieve global consistency. Evaluation on the synthetic Replica and real-world TUM-RGBD, ScanNet, and ScanNet++ datasets demonstrates competitive or superior tracking, mapping, and rendering compared to existing methods for dense RGB-D SLAM. Code is available at loopsplat.github.io.
Liyuan Zhu, Erik Sandström, Konrad Schindler, Iro Armeni
3DV5
2025 Video Depth without Video Models
abstract
Video depth estimation lifts monocular video clips to 3D by inferring dense depth at every frame. Recent advances in single-image depth estimation, brought about by the rise of large foundation models and the use of synthetic training data, have fueled a renewed interest in video depth. However, naively applying a single-image depth estimator to every frame of a video disregards temporal continuity, which not only leads to flickering but may also break when camera motion causes sudden changes in depth range. An obvious and principled solution would be to build on top of video foundation models, but these come with their own limitations; including expensive training and inference, imperfect 3D consistency, and stitching routines for the fixed-length (short) outputs. We take a step back and demonstrate how to turn a single-image latent diffusion model (LDM) into a state-of-the-art video depth estimator. Our model, which we call RollingDepth, has two main ingredients: (i) a multi-frame depth estimator that is derived from a single-image LDM and maps very short video snippets (typically frame triplets) to depth snippets. (ii) a robust, optimization-based registration algorithm that optimally assembles depth snippets sampled at various different frame rates back into a consistent video. RollingDepth is able to efficiently handle long videos with hundreds of frames and delivers more accurate depth videos than both dedicated video depth estimators and high-performing single-frame models. Project page: rollingdepth.github.io.
Bingxin Ke, Dominik Narnhofer, Lei Ke, Torben Peters, Katerina Fragkiadaki, Anton Obukhov, Konrad Schindler
CVPR8
2025 Unsupervised Urban Land Use Mapping with Street View Contrastive Clustering and a Geographical Prior
abstract
Urban land use classification and mapping are critical for urban planning, resource management, and environmental monitoring. Existing remote sensing techniques often lack precision in complex urban environments due to the absence of ground-level details. Unlike aerial perspectives, street view images provide a ground-level view that captures more human and social activities relevant to land use in complex urban scenes. Existing street view-based methods primarily rely on supervised classification, which is challenged by the scarcity of high-quality labeled data and the difficulty of generalizing across diverse urban landscapes. This study introduces an unsupervised contrastive clustering model for street view images with a built-in geographical prior, to enhance clustering performance. When combined with a simple visual assignment of the clusters, our approach offers a flexible and customizable solution to land use mapping, tailored to the specific needs of urban planners. We experimentally show that our method can generate land use maps from geotagged street view image datasets of two cities. As our methodology relies on the universal spatial coherence of geospatial data ("Tobler's law"), it can be adapted to various settings where street view images are available, to enable scalable, unsupervised land use mapping and updating. The code is available at https://github.com/lin102/CCGP.
Lin Che 0001, Yizi Chen, Tanhua Jin, Martin Raubal, Konrad Schindler, Peter Kiefer
SIGSPATIAL/GIS5
2025 Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion
abstract
Depth completion upgrades sparse depth measurements into dense depth maps guided by a conventional image. Existing methods for this highly ill-posed task operate in tightly constrained settings and tend to struggle when applied to images outside the training domain or when the available depth measurements are sparse, irregularly distributed, or of varying density. Inspired by recent advances in monocular depth estimation, we reframe depth completion as an image-conditional depth map generation guided by sparse measurements. Our method, Marigold-DC, builds on a pretrained latent diffusion model for monocular depth estimation and injects the depth observations as test-time guidance via an optimization scheme that runs in tandem with the iterative inference of denoising diffusion. The method exhibits excellent zero-shot generalization across a diverse range of environments and handles even extremely sparse guidance effectively. Our results suggest that contemporary monocular depth priors greatly robustify depth completion: it may be better to view the task as recovering dense depth from (dense) image pixels, guided by sparse depth; rather than as inpainting (sparse) depth, guided by an image. Project website: https://MarigoldDepthCompletion.github.io/
Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker 0002, Konrad Schindler, Anton Obukhov
ICCV6
2025 CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
abstract
We introduce a novel method for generating 360° panoramas from text prompts or images. Our approach leverages recent advances in 3D generation by employing multi-view diffusion models to jointly synthesize the six faces of a cubemap. Unlike previous methods that rely on processing equirectangular projections or autoregressive generation, our method treats each face as a standard perspective image, simplifying the generation process and enabling the use of existing multi-view diffusion models. We demonstrate that these models can be adapted to produce high-quality cubemaps without requiring correspondence-aware attention layers. Our model allows for fine-grained text control, generates high resolution panorama images and generalizes well beyond its training set, whilst achieving state-of-the-art results, both qualitatively and quantitatively.
Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Philipp Henzler, Konrad Schindler, Federico Tombari
ICLR5
2025 GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation
abstract
We propose GALA, a novel representation of 3D shapes that (i) excels at capturing and reproducing complex geometry and surface details, (ii) is computationally efficient, and (iii) lends itself to 3D generative modelling with modern, diffusion-based schemes. The key idea of GALA is to exploit both the global sparsity of surfaces within a 3D volume and their local surface properties. *Sparsity* is promoted by covering only the 3D object boundaries, not empty space, with an ensemble of tree root voxels. Each voxel contains an octree to further limit storage and compute to regions that contain surfaces. *Adaptivity* is achieved by fitting one local and geometry-aware coordinate frame in each non-empty leaf node. Adjusting the orientation of the local grid, as well as the anisotropic scales of its axes, to the local surface shape greatly increases the amount of detail that can be stored in a given amount of memory, which in turn allows for quantization without loss of quality. With our optimized C++/CUDA implementation, GALA can be fitted to an object in less than 10 seconds. Moreover, the representation can efficiently be flattened and manipulated with transformer networks. We provide a cascaded generation pipeline capable of generating 3D shapes with great geometric detail. For more information, please visit our [project page](https://santisy.github.io/GALA/).
Dingdong Yang, Yizhi Wang 0006, Konrad Schindler, Ali Mahdavi-Amiri, Hao (Richard) Zhang
ICLR3
2025 A Variational Perspective on Generative Protein Fitness Optimization
abstract
The goal of protein fitness optimization is to discover new protein variants with enhanced fitness for a given use. The vast search space and the sparsely populated fitness landscape, along with the discrete nature of protein sequences, pose significant challenges when trying to determine the gradient towards configurations with higher fitness. We introduce *Variational Latent Generative Protein Optimization* (VLGPO), a variational perspective on fitness optimization. Our method embeds protein sequences in a continuous latent space to enable efficient sampling from the fitness distribution and combines a (learned) flow matching prior over sequence mutations with a fitness predictor to guide optimization towards sequences with high fitness. VLGPO achieves state-of-the-art results on two different protein benchmarks of varying complexity. Moreover, the variational design with explicit prior and likelihood functions offers a flexible plug-and-play framework that can be easily customized to suit various protein design tasks.
Lea Bogensperger, Dominik Narnhofer, Konrad Schindler, Michael Krauthammer
ICML4
2025 Solving Inverse Problems with FLAIR
abstract
Flow-based latent generative models such as Stable Diffusion 3 are able to generate images with remarkable quality, even enabling photorealistic text-to-image generation. Their impressive performance suggests that these models should also constitute powerful priors for inverse imaging problems, but that approach has not yet led to comparable fidelity. There are several key obstacles: (i) the data likelihood term is usually intractable; (ii) learned generative models cannot be directly conditioned on the distorted observations, leading to conflicting objectives between data likelihood and prior; and (iii) the reconstructions can deviate from the observed data. We present FLAIR, a novel, training-free variational framework that leverages flow-based generative models as prior for inverse problems. To that end, we introduce a variational objective for flow matching that is agnostic to the type of degradation, and combine it with deterministic trajectory adjustments to guide the prior towards regions which are more likely under the posterior. To enforce exact consistency with the observed data, we decouple the optimization of the data fidelity and regularization terms. Moreover, we introduce a time-dependent calibration scheme in which the strength of the regularization is modulated according to off-line accuracy estimates. Results on standard imaging benchmarks demonstrate that FLAIR consistently outperforms existing diffusion- and flow-based methods in terms of reconstruction quality and sample diversity. Source code is available at https://inverseflair.github.io/.
Julius Erbach, Dominik Narnhofer, Andreas Dombos, Bernt Schiele, Jan Eric Lenssen, Konrad Schindler
NeurIPS6
2025 A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking
abstract
The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion (UniVF), a novel and unified framework for video fusion that leverages multi-frame learning and optical flow-based feature warping for informative, temporally coherent video fusion. To support its development, we also introduce Video Fusion Benchmark (VF-Bench), the first comprehensive benchmark covering four video fusion tasks: multi-exposure, multi-focus, infrared-visible, and medical fusion. VF-Bench provides high-quality, well-aligned video pairs obtained through synthetic data generation and rigorous curation from existing datasets, with a unified evaluation protocol that jointly assesses the spatial quality and temporal consistency of video fusion. Extensive experiments show that UniVF achieves state-of-the-art results across all tasks on VF-Bench. Project page: [vfbench.github.io](https://vfbench.github.io).
Zixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui, Lilun Deng, Yulun Zhang 0001, Kai Zhang 0008, Konrad Schindler
NeurIPS8
2025 Fine-Tune Smarter, Not Harder: Parameter-Efficient Fine-Tuning for Geospatial Foundation Models
Francesc Marti Escofet, Benedikt Blumenstiel, Linus Scheibenreif, Paolo Fraccaro, Konrad Schindler
ECML/PKDD (6)5
2025 FlowSDF: Flow Matching for Medical Image Segmentation Using Distance Transforms
abstract
Abstract Medical image segmentation plays an important role in accurately identifying and isolating regions of interest within medical images. Generative approaches are particularly effective in modeling the statistical properties of segmentation masks that are closely related to the respective structures. In this work we introduce FlowSDF, an image-guided conditional flow matching framework, designed to represent the signed distance function (SDF), and, in turn, to represent an implicit distribution of segmentation masks. The advantage of leveraging the SDF is a more natural distortion when compared to that of binary masks. Through the learning of a vector field associated with the probability path of conditional SDF distributions, our framework enables accurate sampling of segmentation masks and the computation of relevant statistical measures. This probabilistic approach also facilitates the generation of uncertainty maps represented by the variance, thereby supporting enhanced robustness in prediction and further analysis. We qualitatively and quantitatively illustrate competitive performance of the proposed method on a public nuclei and gland segmentation data set, highlighting its utility in medical image segmentation applications.
Lea Bogensperger, Dominik Narnhofer, Alexander Falk, Konrad Schindler, Thomas Pock
Int. J. Comput. Vis.4
2024 Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text
abstract
Recently, Transformer-based text detection techniques have sought to predict polygons by encoding the coordinates of individual boundary vertices using distinct query features. However, this approach incurs a significant memory overhead and struggles to effectively capture the intricate relationships between vertices belonging to the same instance. Consequently, irregular text layouts often lead to the prediction of outlined vertices, diminishing the quality of results. To address these challenges, we present an innovative approach rooted in Sparse R-CNN: a cascade decoding pipeline for polygon prediction. Our method ensures precision by iteratively refining polygon predictions, considering both the scale and location of preceding results. Leveraging this stabilized regression pipeline, even employing just a single feature vector to guide polygon instance regression yields promising detection results. Simultaneously, the leverage of instance-level feature proposal substantially enhances memory efficiency ( > 50% less vs. the SOTA method DPText-DETR) and reduces inference speed (> 40% less vs. DPText-DETR) with comparable performance on benchmarks. The code is available at https://github.com/Albertchen98/Box2Poly.git.
Konrad Schindler, Nicolò Savioli, Liqiu Meng
AAAI3
2024 Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
abstract
Monocular depth estimation is a fundamental computer vision task. Recovering 3D depth from a single image is geometrically ill-posed and requires scene understanding, so it is not surprising that the rise of deep learning has led to a breakthrough. The impressive progress of monocular depth estimators has mirrored the growth in model capacity, from relatively modest CNNs to large Transformer architectures. Still, monocular depth estimators tend to struggle when presented with images with unfamiliar content and layout, since their knowledge of the visual world is restricted by the data seen during training, and challenged by zero-shot generalization to new domains. This motivates us to explore whether the extensive priors captured in recent generative diffusion models can enable better, more generalizable depth estimation. We introduce Marigold, a method for affine-invariant monocular depth estimation that is derived from Stable Diffusion and retains its rich prior knowledge. The estimator can be fine-tuned in a couple of days on a single GPU using only synthetic training data. It delivers state-of-the-art performance across a wide range of datasets, including over 20% performance gains in specific cases. Project page: https://marigoldmonodepth.github.io.
Bingxin Ke, Anton Obukhov, Nando Metzger, Rodrigo Caye Daudt, Konrad Schindler
CVPR6
2024 Point2CAD: Reverse Engineering CAD Models from 3D Point Clouds
abstract
Computer-Aided Design (CAD) model reconstruction from point clouds is an important problem at the intersection of computer vision, graphics, and machine learning; it saves the designer significant time when iterating on in-the-wild objects. Recent advancements in this direction achieve relatively reliable semantic segmentation but still struggle to produce an adequate topology of the CAD model. In this work, we analyze the current state of the art for that ill-posed task and identify shortcomings of existing methods. We propose a hybrid analyticneural reconstruction scheme that bridges the gap between segmented point clouds and structured CAD models and can be readily combined with different segmentation backbones. Moreover, to power the surface fitting stage, we propose a novel implicit neural representation of freeform surfaces, driving up the performance of our overall CAD reconstruction scheme. We extensively evaluate our method on the popular ABC benchmark of CAD models and set a new state-of-the-art for that dataset. Code is available at https://github.com/YujiaLiu76/point2cad.
Yujia Liu 0002, Anton Obukhov, Jan Dirk Wegner, Konrad Schindler
CVPR4
2024 StegoGAN: Leveraging Steganography for Non-Bijective Image-to-Image Translation
abstract
Most image-to-image translation models postulate that a unique correspondence exists between the semantic classes of the source and target domains. However, this assumption does not always hold in real-world scenarios due to divergent distributions, different class sets, and asymmet- rical information representation. As conventional GANs attempt to generate images that match the distribution of the target domain, they may hallucinate spurious instances of classes absent from the source domain, thereby dimin- ishing the usefulness and reliability of translated images. CycleGAN-based methods are also known to hide the mis- matched information in the generated images to bypass cy- cle consistency objectives, a process known as steganogra- phy. In response to the challenge of non-bijective image translation, we introduce StegoGAN, a novel model that leverages steganography to prevent spurious features in generated images. Our approach enhances the semantic consistency of the translated images without requiring ad- ditional postprocessing or supervision. Our experimental evaluations demonstrate that StegoGAN outperforms existing GAN-based models across various non-bijective image- to-image translation tasks, both qualitatively and quantita- tively. Our code and pretrained models are accessible at https://github.com/sian-wusidi/StegoGAN.
Sidi Wu 0001, Yizi Chen, Samuel Mermet, Lorenz Hurni, Konrad Schindler, Nicolas Gonthier, Loïc Landrieu
CVPR5
2024 Dynamic LiDAR Re-Simulation Using Compositional Neural Fields
abstract
We introduce DyNFL, a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments, accompanied by bounding boxes of moving objects, to construct an editable neural field. This field, comprising separately reconstructed static background and dynamic objects, allows users to modify viewpoints, adjust object positions, and seamlessly add or remove objects in the re-simulated scene. A key innovation of our method is the neural field composition technique, which effectively integrates reconstructed neural assets from various scenes through a ray drop test, accounting for occlusions and transparent surfaces. Our evaluation with both synthetic and real-world environments demonstrates that DyNFL substantially improves dynamic scene LiDAR simulation, offering a combination of physical fidelity and flexible editing capabilities. [project page]
Hanfeng Wu, Xingxing Zuo 0001, Stefan Leutenegger, Or Litany, Konrad Schindler
CVPR5
2024 Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D Environments
abstract
Research into dynamic 3D scene understanding has pri-marily focused on short-term change tracking from dense observations, while little attention has been paid to long-term changes with sparse observations. We address this gap with More2,a novel approach for multi-object relo-calization and reconstruction in evolving environments. We view these environments as “living scenes” and consider the problem of transforming scans taken at different points in time into a 3D reconstruction of the object instances, whose accuracy and completeness increase over time. At the core of our method lies an SE (3)-equivariant represen-tation in a single encoder-decoder network, trained on syn-thetic data. This representation enables us to seamlessly tackle instance matching, registration, and reconstruction. We also introduce a joint optimization algorithm that facil-itates the accumulation of point clouds originating from the same instance across multiple scans taken at different points in time. We validate our method on synthetic and real-world data and demonstrate state-of-the-art performance in both end-to-end performance and individual subtasks. [project]
Liyuan Zhu, Konrad Schindler, Iro Armeni
CVPR3
2024 DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic Control
Yuru Jia, Lukas Hoyer, Tianfu Wang 0006, Luc Van Gool, Konrad Schindler, Anton Obukhov
ECCV (39)6
2024 TetraDiffusion: Tetrahedral Diffusion Models for 3D Shape Generation
Nikolai Kalischek, Torben Peters, Jan Dirk Wegner, Konrad Schindler
ECCV (53)4
2024 AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation
abstract
During interactive segmentation, a model and a user work together to delineate objects of interest in a 3D point cloud. In an iterative process, the model assigns each data point to an object (or the background), while the user corrects errors in the resulting segmentation and feeds them back into the model. The current best practice formulates the problem as binary classification and segments objects one at a time. The model expects the user to provide positive clicks to indicate regions wrongly assigned to the background and negative clicks on regions wrongly assigned to the object. Sequentially visiting objects is wasteful since it disregards synergies between objects: a positive click for a given object can, by definition, serve as a negative click for nearby objects. Moreover, a direct competition between adjacent objects can speed up the identification of their common boundary. We introduce AGILE3D, an efficient, attention-based model that (1) supports simultaneous segmentation of multiple 3D objects, (2) yields more accurate segmentation masks with fewer user clicks, and (3) offers faster inference. Our core idea is to encode user clicks as spatial-temporal queries and enable explicit interactions between click queries as well as between them and the 3D scene through a click attention module. Every time new clicks are added, we only need to run a lightweight decoder that produces updated segmentation masks. In experiments with four different 3D point cloud datasets, AGILE3D sets a new state-of-the-art. Moreover, we also verify its practicality in real-world setups with real user studies. Project page: https://ywyue.github.io/AGILE3D.
Yuanwen Yue, Sabarinath Mahadevan, Jonas Schult, Francis Engelmann, Bastian Leibe, Konrad Schindler, Theodora Kontogianni
ICLR6
2024 Modelling the Troposphere with Global Navigation Satellite Systems, Meteorological Data and Machine Learning
abstract
Global Navigation Satellite Systems (GNSS), such as the American Global Positioning System (GPS) and the European Galileo system, are capable of monitoring tropospheric properties. An important parameter describing the tropospheric impact on GNSS is zenith wet delay (ZWD), which is highly correlated to the amount of water vapour in the troposphere and thus interesting for atmospheric and climate research. This work demonstrates how GNSS observations help to sense the atmosphere and its dynamics by using a newly developed machine learning-based ZWD model. The model provides ZWD globally for the years 2010 to 2023 with a positive trend in the Northern Hemisphere and a negative trend in the Southern Hemisphere. Furthermore, the global average ZWD anomaly follows alternating trends, strongly correlated with the El Niño Southern Oscillation (ENSO) index, increasing up to a correlation coefficient of 0.74 when introducing a time lag of two months.
Laura Crocetti, Matthias Schartner, Konrad Schindler, Rochelle Schneider, Benedikt Soja
IGARSS3
2024 Estimating Fine-Grained Population Growth Rates from Coarse Census Data
abstract
Fine-grained population estimation is important for several domains, such as urban planning, public health, and humanitarian action. Due to limited resources, population maps with sufficient spatial resolution and temporal frequency are not available for many developing countries. Population estimates are often based on statistics available at the national or provincial level, which moreover are updated infrequently. The United Nations produce estimates of population growth rates, which are widely used to project population numbers to a target year, starting from the most recent census data. However, they use a simplified model of population growth with uniform growth rates for all urban, respectively rural areas within a country. This neglects the complex dynamics of population growth (e.g., growth rates in big cities are usually larger than in smaller urban areas) and leads to significant errors in population projections. In this work, we propose a methodology to estimate fine-grained population growth rates and present experimental results for Mozambique.
John E. Vargas-Munoz, Nando Metzger, Rodrigo Caye Daudt, Konrad Schindler, Devis Tuia
IGARSS4
2024 BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation
abstract
By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically complex scenes that challenge their geometry prior, trained on less diverse 3D data. To leverage the complementary merits of both worlds, we propose BetterDepth to achieve geometrically correct affine-invariant MDE while capturing fine details. Specifically, BetterDepth is a conditional diffusion-based refiner that takes the prediction from pre-trained MDE models as depth conditioning, in which the global depth layout is well-captured, and iteratively refines details based on the input image. For the training of such a refiner, we propose global pre-alignment and local patch masking methods to ensure BetterDepth remains faithful to the depth conditioning while learning to add fine-grained scene details. With efficient training on small-scale synthetic datasets, BetterDepth achieves state-of-the-art zero-shot MDE performance on diverse public datasets and on in-the-wild scenes. Moreover, BetterDepth can improve the performance of other MDE models in a plug-and-play manner without further re-training.
Xiang Zhang 0022, Bingxin Ke, Hayko Riemenschneider, Nando Metzger, Anton Obukhov, Markus Gross 0001, Konrad Schindler, Christopher Schroers
NeurIPS7
2024 Recognition of Unseen Bird Species by Learning from Field Guides
abstract
We exploit field guides to learn bird species recognition, in particular zero-shot recognition of unseen species. Illustrations contained in field guides deliberately focus on discriminative properties of each species, and can serve as side information to transfer knowledge from seen to unseen bird species. We study two approaches: (1) a contrastive encoding of illustrations, which can be fed into standard zero-shot learning schemes; and (2) a novel method that leverages the fact that illustrations are also images and as such structurally more similar to photographs than other kinds of side information. Our results show that illustrations from field guides, which are readily available for a wide range of species, are indeed a competitive source of side information for zero-shot learning. On a subset of the iNaturalist2021 dataset with 749 seen and 739 unseen species, we obtain a classification accuracy of unseen bird species of 12% @top-1 and 38% @top-10, which shows the potential of field guides for challenging real-world scenarios with many species. Our code is available at https://github.com/ac-rodriguez/zsl_billow.
Andrés C. Rodríguez, Stefano D'Aronco, Rodrigo Caye Daudt, Jan Dirk Wegner, Konrad Schindler
WACV5
2023 Breathing New Life into 3D Assets with Generative Repainting
Tianfu Wang 0006, Menelaos Kanakis, Konrad Schindler, Luc Van Gool, Anton Obukhov
BMVC3
2023 BiasBed - Rigorous Texture Bias Evaluation
abstract
The well-documented presence of texture bias in modern convolutional neural networks has led to a plethora of algorithms that promote an emphasis on shape cues, often to support generalization to new domains. Yet, common datasets, benchmarks and general model selection strategies are missing, and there is no agreed, rigorous evaluation protocol. In this paper, we investigate difficulties and limitations when training networks with reduced texture bias. In particular, we also show that proper evaluation and meaningful comparisons between methods are not trivial. We introduce BiasBed, a testbed for texture- and style-biased training, including multiple datasets and a range of existing algorithms. It comes with an extensive evaluation protocol that includes rigorous hypothesis testing to gauge the significance of the results, despite the considerable training instability of some style bias methods. Our extensive experiments, shed new light on the need for careful, statistically founded evaluation protocols for style bias (and beyond). E.g., we find that some algorithms proposed in the literature do not significantly mitigate the impact of style bias at all. With the release of BiasBed, we hope to foster a common understanding of consistent and meaningful comparisons, and consequently faster progress towards learning methods free of texture bias. Code is available at https://github.com/D1noFuzi/BiasBed
Nikolai Kalischek, Rodrigo Caye Daudt, Torben Peters, Reinhard Furrer, Jan Dirk Wegner, Konrad Schindler
CVPR6
2023 Guided Depth Super-Resolution by Deep Anisotropic Diffusion
abstract
Performing super-resolution of a depth image using the guidance from an RGB image is a problem that concerns several fields, such as robotics, medical imaging, and remote sensing. While deep learning methods have achieved good results in this problem, recent work highlighted the value of combining modern methods with more formal frameworks. In this work, we propose a novel approach which combines guided anisotropic diffusion with a deep convolutional network and advances the state of the art for guided depth super-resolution. The edge transferring/enhancing properties of the diffusion are boosted by the contextual reasoning capabilities of modern networks, and a strict adjustment step guarantees perfect adherence to the source image. We achieve unprecedented results in three commonly used benchmarks for guided depth superresolution. The performance gain compared to other methods is the largest at larger scales, such as × 32 scaling. Code11https://github.com/prs-eth/Diffusion-Super-Resolution for the proposed method is available to promote reproducibility of our results.
Nando Metzger, Rodrigo Caye Daudt, Konrad Schindler
CVPR3
2023 BITE: Beyond Priors for Improved Three-D Dog Pose Estimation
abstract
We address the problem of inferring the 3D shape and pose of dogs from images. Given the lack of 3D training data, this problem is challenging, and the best methods lag behind those designed to estimate human shape and pose. To make progress, we attack the problem from multiple sides at once. First, we need a good 3D shape prior, like those available for humans. To that end, we learn a dog-specific 3D parametric model, called D-SMAL. Second, existing methods focus on dogs in standing poses because when they sit or lie down, their legs are self occluded and their bodies deform. Without access to a good pose prior or 3D data, we need an alternative approach. To that end, we exploit contact with the ground as a form of side information. We consider an existing large dataset of dog images and label any 3D contact of the dog with the ground. We exploit body-ground contact in estimating dog pose and find that it significantly improves results. Third, we develop a novel neural network architecture to infer and exploit this contact information. Fourth, to make progress, we have to be able to measure it. Current evaluation metrics are based on 2D features like keypoints and silhouettes, which do not directly correlate with 3D errors. To address this, we create a synthetic dataset containing rendered images of scanned 3D dogs. With these advances, our method recovers significantly better dog shape and pose than the state of the art, and we evaluate this improvement in 3D. Our code, model and test dataset are publicly available for research purposes at https://bite.is.tue.mpg.de/.
Nadine Bertsch, Shashank Tripathi, Konrad Schindler, Michael J. Black, Silvia Zuffi
CVPR3
2023 Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries
abstract
We address 2D floorplan reconstruction from 3D scans. Existing approaches typically employ heuristically designed multi-stage pipelines. Instead, we formulate floor-plan reconstruction as a single-stage structured prediction task: find a variablesize set of polygons, which in turn are variable-length sequences of ordered vertices. To solve it we develop a novel Transformer architecture that generates polygons of multiple rooms in parallel, in a holistic manner without hand-crafted intermediate stages. The model features two-level queries for polygons and corners, and includes polygon matching to make the network end-to-end trainable. Our method achieves a new state-of-the-art for two challenging datasets, Structured3D and SceneCAD, along with significantly faster inference than previous methods. Moreover, it can readily be extended to predict additional information, i.e., semantic room types and architectural elements like doors and windows. Our code and models are available at: https://github.com/ywyue/RoomFormer.
Yuanwen Yue, Theodora Kontogianni, Konrad Schindler, Francis Engelmann
CVPR3
2023 Cross-attention Spatio-temporal Context Transformer for Semantic Segmentation of Historical Maps
abstract
Historical maps provide useful spatio-temporal information on the Earth's surface before modern earth observation techniques came into being. To extract information from maps, neural networks, which gain wide popularity in recent years, have replaced hand-crafted map processing methods and tedious manual labor. However, aleatoric uncertainty, known as data-dependent uncertainty, inherent in the drawing/scanning/fading defects of the original map sheets and inadequate contexts when cropping maps into small tiles considering the memory limits of the training process, challenges the model to make correct predictions. As aleatoric uncertainty cannot be reduced even with more training data collected, we argue that complementary spatio-temporal contexts can be helpful. To achieve this, we propose a U-Net-based network that fuses spatio-temporal features with cross-attention transformers (U-SpaTem), aggregating information at a larger spatial range as well as through a temporal sequence of images. Our model achieves a better performance than other state-or-art models that use either temporal or spatial contexts. Compared with pure vision transformers, our model is more lightweight and effective. To the best of our knowledge, leveraging both spatial and temporal contexts have been rarely explored before in the segmentation task. Even though our application is on segmenting historical maps, we believe that the method can be transferred into other fields with similar problems like temporal sequences of satellite images. Our code is freely accessible at https://github.com/chenyizi086/wu.2023.sigspatial.git.
Sidi Wu 0001, Yizi Chen, Konrad Schindler, Lorenz Hurni
SIGSPATIAL/GIS3
2023 Neural LiDAR Fields for Novel View Synthesis
abstract
We present Neural Fields for LiDAR (NFL), a method to optimise a neural field scene representation from LiDAR measurements, with the goal of synthesizing realistic LiDAR scans from novel viewpoints. NFL combines the rendering power of neural fields with a detailed, physically motivated model of the LiDAR sensing process, thus enabling it to accurately reproduce key sensor behaviors like beam divergence, secondary returns, and ray dropping. We evaluate NFL on synthetic and real LiDAR scans and show that it outperforms explicit reconstruct-then-simulate methods as well as other NeRF-style methods on LiDAR novel view synthesis task. Moreover, we show that the improved realism of the synthesized views narrows the domain gap to real scans and translates to better registration and semantic segmentation performance.
Zan Gojcic, Francis Williams, Yoni Kasten, Sanja Fidler, Konrad Schindler, Or Litany
ICCV7
2023 Interactive Object Segmentation in 3D Point Clouds
abstract
We propose an interactive approach for 3D instance segmentation, where users can iteratively collaborate with a deep learning model to segment objects directly in a 3D point cloud. Current methods for 3D instance segmentation are generally trained in a fully-supervised fashion, which requires large amounts of costly training labels, and does not generalize well to classes unseen during training. Few works have attempted to obtain 3D segmentation masks using human interactions. Existing methods rely on user feedback in the 2D image domain. As a consequence, users are required to constantly switch between 2D images and 3D representations, and custom architectures are employed to combine multiple input modalities. Therefore, integration with existing standard 3D models is not straightforward. The core idea of this work is to enable users to interact directly with 3D point clouds by clicking on desired 3D objects of interest (or their background) to interactively segment the scene in an open-world setting. Specifically, our method does not require training data from any target domain and can adapt to new environments where no appropriate training sets are available. Our system continuously adjusts the object segmentation based on the user feedback and achieves accurate dense 3D segmentation masks with minimal human effort (few clicks per object). Besides its potential for efficient labeling of large-scale and varied 3D datasets, our approach, where the user directly interacts with the 3D environment, enables new AR/VR and human-robot interaction applications.
Theodora Kontogianni, Ekin Celikkan, Siyu Tang 0001, Konrad Schindler
ICRA4
2023 BARC: Breed-Augmented Regression Using Classification for 3D Dog Reconstruction from Images
abstract
Abstract The goal of this work is to reconstruct 3D dogs from monocular images. We take a model-based approach, where we estimate the shape and pose parameters of a 3D articulated shape model for dogs. We consider dogs as they constitute a challenging problem, given they are highly articulated and come in a variety of shapes and appearances. Recent work has considered a similar task using the multi-animal SMAL model, with additional limb scale parameters, obtaining reconstructions that are limited in terms of realism. Like previous work, we observe that the original SMAL model is not expressive enough to represent dogs of many different breeds. Moreover, we make the hypothesis that the supervision signal used to train the network, that is 2D keypoints and silhouettes, is not sufficient to learn a regressor that can distinguish between the large variety of dog breeds. We therefore go beyond previous work in two important ways. First, we modify the SMAL shape space to be more appropriate for representing dog shape. Second, we formulate novel losses that exploit information about dog breeds. In particular, we exploit the fact that dogs of the same breed have similar body shapes. We formulate a novel breed similarity loss, consisting of two parts: One term is a triplet loss, that encourages the shape of dogs from the same breed to be more similar than dogs of different breeds. The second one is a breed classification loss. With our approach we obtain 3D dogs that, compared to previous work, are quantitatively better in terms of 2D reconstruction, and significantly better according to subjective and quantitative 3D evaluations. Our work shows that a-priori side information about similarity of shape and appearance, as provided by breed labels, can help to compensate for the lack of 3D training data. This concept may be applicable to other animal species or groups of species. We call our method BARC (Breed-Augmented Regression using Classification). Our code is publicly available for research purposes at https://barc.is.tue.mpg.de/ .
Nadine Bertsch, Silvia Zuffi, Konrad Schindler, Michael J. Black
Int. J. Comput. Vis.3
2023 Fine-Grained Species Recognition With Privileged Pooling: Better Sample Efficiency Through Supervised Attention
abstract
We propose a scheme for supervised image classification that uses privileged information, in the form of keypoint annotations for the training data, to learn strong models from small and/or biased training sets. Our main motivation is the recognition of animal species for ecological applications such as biodiversity modelling, which is challenging because of long-tailed species distributions due to rare species, and strong dataset biases such as repetitive scene background in camera traps. To counteract these challenges, we propose a visual attention mechanism that is supervised via keypoint annotations that highlight important object parts. This privileged information, implemented as a novel privileged pooling operation, is only required during training and helps the model to focus on regions that are discriminative. In experiments with three different animal species datasets, we show that deep networks with privileged pooling can use small training sets more efficiently and generalize better.
Andrés C. Rodríguez, Stefano D'Aronco, Konrad Schindler, Jan Dirk Wegner
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 U-TILISE: A Sequence-to-Sequence Model for Cloud Removal in Optical Satellite Time Series
abstract
Satellite image time series in the optical and infrared spectrum suffer from frequent data gaps due to cloud cover, cloud shadows, and temporary sensor outages. It has been a long-standing problem of remote sensing research how to best reconstruct the missing pixel values and obtain complete, cloud-free image sequences. We approach that problem from the perspective of representation learning and develop U-TILISE, an efficient neural model that is able to implicitly capture spatio-temporal patterns of the spectral intensities, and that can therefore be trained to map a cloud-masked input sequence to a cloud-free output sequence. The model consists of a convolutionalspatial encoderthat maps each individual frame of the input sequence to a latent encoding; an attention-basedtemporal encoderthat captures dependencies between those per-frame encodings and lets them exchange information along the time dimension; and a convolutionalspatial decoderthat decodes the latent embeddings back into multi-spectral images. We experimentally evaluate the proposed model on EarthNet2021, a dataset of Sentinel-2 time series acquired all over Europe, and demonstrate its superior ability to reconstruct the missing pixels. Compared to a standard interpolation baseline, it increases the PSNR by 1.8 dB at previously seen locations and by 1.3 dB at unseen locations.
Corinne Stucker, Vivien Sainte Fare Garnot, Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.3
2022 Message from the Program Chairs: 3DV 2022
abstract
We welcome you to the 2022 edition of the International Conference on 3D Vision (3DV 2022). The conference took place in hybrid format: after the hiatus due to the global pandemic, we are happy to be able to host an in-person meeting in Prague, Czech Republic, while attendees who could not travel to Prague still also had the possibility to participate virtually.
Angela Dai, Jana Kosecka, Gin Hee Lee, Konrad Schindler
3DV4
2022 T4DT: Tensorizing Time for Learning Temporal 3D Visual Data
Mikhail Usvyatsov, Rafael Ballester-Ripoll, Lina Bashaeva, Konrad Schindler, Gonzalo Ferrer 0001, Ivan V. Oseledets
BMVC4
2022 Learning Graph Regularisation for Guided Super-Resolution
abstract
We introduce a novel formulation for guided super-resolution. Its core is a differentiable optimisation layer that operates on a learned affinity graph. The learned graph potentials make it possible to leverage rich contextual information from the guide image, while the explicit graph optimisation within the architecture guarantees rigorous fidelity of the high-resolution target to the low-resolution source. With the decision to employ the source as a constraint rather than only as an input to the prediction, our method differs from state-of-the-art deep architectures for guided super-resolution, which produce targets that, when downsampled, will only approximately reproduce the source. This is not only theoretically appealing, but also produces crisper, more natural-looking images. A key property of our method is that, although the graph connectivity is restricted to the pixel lattice, the associated edge potentials are learned with a deep feature extractor and can encode rich context information over large receptive fields. By taking advantage of the sparse graph connectivity, it becomes possible to propagate gradients through the optimisation layer and learn the edge potentials from data. We extensively evaluate our method on several datasets, and consistently outperform recent baselines in terms of quantitative reconstruction errors, while also delivering visually sharper outputs. Moreover, we demonstrate that our method generalises particularly well to new datasets not seen during training.
Riccardo de Lutio, Alexander Becker 0002, Stefano D'Aronco, Stefania Russo, Jan Dirk Wegner, Konrad Schindler
CVPR6
2022 BARC: Learning to Regress 3D Dog Shape from Images by Exploiting Breed Information
abstract
Our goal is to recover the 3D shape and pose of dogs from a single image. This is a challenging task because dogs exhibit a wide range of shapes and appearances, and are highly articulated. Recent work has proposed to directly regress the SMAL animal model, with additional limb scale parameters, from images. Our method, called BARC (Breed-Augmented Regression using Classification), goes beyond prior work in several important ways. First, we modify the SMAL shape space to be more appropriate for representing dog shape. But, even with a better shape model, the problem of regressing dog shape from an image is still challenging because we lack paired images with 3D ground truth. To compensate for the lack of paired data, we formulate novel losses that exploit information about dog breeds. In particular, we exploit the fact that dogs of the same breed have similar body shapes. We formulate a novel breed similarity loss consisting of two parts: One term encourages the shape of dogs from the same breed to be more similar than dogs of different breeds. The second one, a breed classification loss, helps to produce recognizable breed-specific shapes. Through ablation studies, we find that our breed losses significantly improve shape accuracy over a baseline without them. We also compare BARC qualitatively to WLDO with a perceptual study and find that our approach produces dogs that are significantly more realistic. This work shows that a-priori information about genetic similarity can help to compensate for the lack of 3D training data. This concept may be applicable to other animal species or groups of species. Our code is publicly available for research purposes at https://barc.is.tue.mpg.de/.
Nadine Bertsch, Silvia Zuffi, Konrad Schindler, Michael J. Black
CVPR3
2022 Dynamic 3D Scene Analysis by Point Cloud Accumulation
Zan Gojcic, Andreas Wieser, Konrad Schindler
ECCV (38)5
2022 FiLM-Ensemble: Probabilistic Deep Learning via Feature-wise Linear Modulation
abstract
The ability to estimate epistemic uncertainty is often crucial when deploying machine learning in the real world, but modern methods often produce overconfident, uncalibrated uncertainty predictions. A common approach to quantify epistemic uncertainty, usable across a wide class of prediction models, is to train a model ensemble. In a naive implementation, the ensemble approach has high computational cost and high memory demand. This challenges in particular modern deep learning, where even a single deep network is already demanding in terms of compute and memory, and has given rise to a number of attempts to emulate the model ensemble without actually instantiating separate ensemble members. We introduce FiLM-Ensemble, a deep, implicit ensemble method based on the concept of Feature-wise Linear Modulation (FiLM). That technique was originally developed for multi-task learning, with the aim of decoupling different tasks. We show that the idea can be extended to uncertainty quantification: by modulating the network activations of a single deep network with FiLM, one obtains a model ensemble with high diversity, and consequently well-calibrated estimates of epistemic uncertainty, with low computational overhead in comparison. Empirically, FiLM-Ensemble outperforms other implicit ensemble methods, and it comes very close to the upper bound of an explicit ensemble of networks (sometimes even beating it), at a fraction of the memory cost.
Mehmet Ozgur Turkoglu, Alexander Becker 0002, Hüseyin Anil Gündüz, Mina Rezaei, Bernd Bischl, Rodrigo Caye Daudt, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler
NeurIPS9
2022 tntorch: Tensor Network Learning with PyTorch
abstract
We present tntorch, a tensor learning framework that supports multiple decompositions (including Candecomp/Parafac, Tucker, and Tensor Train) under a unified interface. With our library, the user can learn and handle low-rank tensors with automatic differentiation, seamless GPU support, and the convenience of PyTorch's API. Besides decomposition algorithms, tntorch implements differentiable tensor algebra, rank truncation, cross-approximation, batch processing, comprehensive tensor arithmetics, and more.
Mikhail Usvyatsov, Rafael Ballester-Ripoll, Konrad Schindler
J. Mach. Learn. Res.3
2022 Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer
abstract
The success of monocular depth estimation relies on large and diverse training sets. Due to the challenges associated with acquiring dense ground-truth depth across different environments at scale, a number of datasets with distinct characteristics and biases have emerged. We develop tools that enable mixing multiple datasets during training, even if their annotations are incompatible. In particular, we propose a robust training objective that is invariant to changes in depth range and scale, advocate the use of principled multi-objective learning to combine data from different sources, and highlight the importance of pretraining encoders on auxiliary tasks. Armed with these tools, we experiment with five diverse training datasets, including a new, massive data source: 3D films. To demonstrate the generalization power of our approach we use zero-shot cross-dataset transfer, i.e. we evaluate on datasets that were not seen during training. The experiments confirm that mixing data from complementary sources greatly improves monocular depth estimation. Our approach clearly outperforms competing methods across diverse datasets, setting a new state of the art for monocular depth estimation.
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, Vladlen Koltun
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Gating Revisited: Deep Multi-Layer RNNs That can be Trained
abstract
We propose a new STAckable Recurrent cell (STAR) for recurrent neural networks (RNNs), which has fewer parameters than widely used LSTM [16] and GRU [10] while being more robust against vanishing or exploding gradients. Stacking recurrent units into deep architectures suffers from two major limitations: (i) many recurrent cells (e.g., LSTMs) are costly in terms of parameters and computation resources; and (ii) deep RNNs are prone to vanishing or exploding gradients during training. We investigate the training of multi-layer RNNs and examine the magnitude of the gradients as they propagate through the network in the "vertical" direction. We show that, depending on the structure of the basic recurrent unit, the gradients are systematically attenuated or amplified. Based on our analysis we design a new type of gated cell that better preserves gradient magnitude. We validate our design on a large number of sequence modelling tasks and demonstrate that the proposed STAR cell allows to build and train deeper recurrent architectures, ultimately leading to improved performance while being computationally more efficient.
Mehmet Ozgur Turkoglu, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Crop Classification Under Varying Cloud Cover With Neural Ordinary Differential Equations
abstract
Optical satellite sensors cannot see the earth’s surface through clouds. Despite the periodic revisit cycle, image sequences acquired by earth observation satellites are, therefore,irregularlysampled in time. State-of-the-art methods for crop classification (and other time-series analysis tasks) rely on techniques that implicitly assume regular temporal spacing between observations, such as recurrent neural networks (RNNs). We propose to use neural ordinary differential equations (NODEs) in combination with RNNs to classify crop types in irregularly spaced image sequences. The resulting ODE-RNN models consist of two steps: an update step, where a recurrent unit assimilates new input data into the model’s hidden state, and a prediction step, in which NODE propagates the hidden state until the next observation arrives. The prediction step is based on a continuous representation of the latent dynamics, which has several advantages. At the conceptual level, it is a more natural way to describe the mechanisms that govern the phenological cycle. From a practical point of view, it makes it possible to sample the system state at arbitrary points in time such that one can integrate observations whenever they are available and extrapolate beyond the last observation. Our experiments show that ODE-RNN, indeed, improves classification accuracy over common baselines, such as LSTM, GRU, temporal convolutional network, and transformer. The gains are most prominent in the challenging scenario where only few observations are available (i.e., frequent cloud cover). Moreover, we show that the ability to extrapolate translates to better classification performance early in the season, which is important for forecasting.
Nando Metzger, Mehmet Ozgur Turkoglu, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.5
2022 Learning a Joint Embedding of Multiple Satellite Sensors: A Case Study for Lake Ice Monitoring
abstract
Fusing satellite imagery acquired with different sensors has been a long-standing challenge of Earth observation, particularly across different modalities such as optical and Synthetic Aperture Radar (SAR) images. Here, we explore the joint analysis of imagery from different sensors in the light of representation learning: we propose to learn a joint embedding of multiple satellite sensors within a deep neural network. Our application problem is the monitoring of lake ice on Alpine lakes. To reach the temporal resolution requirement of the Swiss Global Climate Observing System (GCOS) office, we combine three image sources: Sentinel-1 SAR (S1-SAR), Terra MODIS and Suomi-NPP VIIRS. The large gaps between the optical and SAR domains and between the sensor resolutions make this a challenging instance of the sensor fusion problem. Our approach can be classified as a late fusion that is learnt in a data-driven manner. The proposed network architecture has separate encoding branches for each image sensor, which feed into a single latent embedding. I.e., a common feature representation shared by all inputs, such that subsequent processing steps deliver comparable output irrespective of which sort of input image was used. By fusing satellite data, we map lake ice at a temporal resolution of91% (respectively, mIoU scores >60%) and generalises well across different lakes and winters. Moreover, it sets a new state-of-the-art for determining the important ice-on and ice-off dates for the target lakes, in many cases meeting the GCOS requirement.
Manu Tom, Yuchang Jiang, Emmanuel Baltsavias, Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.4
2021 Visual Camera Re-Localization Using Graph Neural Networks and Relative Pose Supervision
abstract
Visual re-localization means using a single image as input to estimate the camera’s location and orientation relative to a pre-recorded environment. The highest-scoring methods are “structure-based,” and need the query camera’s intrinsics as an input to the model, with careful geometric optimization. When intrinsics are absent, methods vie for accuracy by making various other assumptions. This yields fairly good localization scores, but the models are “narrow” in some way, e.g., requiring costly test-time computations, or depth sensors, or multiple query frames. In contrast, our proposed method makes few special assumptions, and is fairly lightweight in training and testing.Our pose regression network learns from only relative poses of training scenes. For inference, it builds a graph connecting the query image to training counterparts and uses a graph neural network (GNN) with image representations on nodes and image-pair representations on edges. By efficiently passing messages between them, both representation types are refined to produce a consistent camera pose estimate. We validate the effectiveness of our approach on both standard indoor (7-Scenes) and outdoor (Cambridge Landmarks) camera re-localization benchmarks. Our relative pose regression method matches the accuracy of absolute pose regression networks, while retaining the relative-pose models’ test-time speed and ability to generalize to non-training scenes.
Mehmet Ozgur Turkoglu, Eric Brachmann, Konrad Schindler, Gabriel J. Brostow, Áron Monszpart
3DV3
2021 Predator: Registration of 3D Point Clouds With Low Overlap
abstract
We introduce PREDATOR, a model for pairwise point-cloud registration with deep attention to the overlap region. Different from previous work, our model is specifically designed to handle (also) point-cloud pairs with low overlap. Its key novelty is an overlap-attention block for early information exchange between the latent encodings of the two point clouds. In this way the subsequent decoding of the latent representations into per-point features is conditioned on the respective other point cloud, and thus can predict which points are not only salient, but also lie in the overlap region between the two point clouds. The ability to focus on points that are relevant for matching greatly improves performance: PREDATOR raises the rate of successful registrations by more than 20% in the low-overlap scenario, and also sets a new state of the art for the 3DMatch benchmark with 89% registration recall. [Code release]
Zan Gojcic, Mikhail Usvyatsov, Andreas Wieser, Konrad Schindler
CVPR5
2021 In the Light of Feature Distributions: Moment Matching for Neural Style Transfer
abstract
Style transfer aims to render the content of a given image in the graphical/artistic style of another image. The fundamental concept underlying Neural Style Transfer (NST) is to interpret style as a distribution in the feature space of a Convolutional Neural Network, such that a desired style can be achieved by matching its feature distribution. We show that most current implementations of that concept have important theoretical and practical limitations, as they only partially align the feature distributions. We propose a novel approach that matches the distributions more precisely, thus reproducing the desired style more faithfully, while still being computationally efficient. Specifically, we adapt the dual form of Central Moment Discrepancy (CMD), as recently proposed for domain adaptation, to minimize the difference between the target style and the feature distribution of the output image. The dual interpretation of this metric explicitly matches all higher-order centralized moments and is therefore a natural extension of existing NST methods that only take into account the first and second moments. Our experiments confirm that the strong theoretical properties also translate to visually better style transfer, and better disentangle style from semantic image content.
Nikolai Kalischek, Jan Dirk Wegner, Konrad Schindler
CVPR3
2021 Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation
abstract
We propose an end-to-end trainable framework that processes large-scale visual data tensors by looking at a fraction of their entries only. Our method combines a neural network encoder with a tensor train decomposition to learn a low-rank latent encoding, coupled with cross-approximation (CA) to learn the representation through a subset of the original samples. CA is an adaptive sampling algorithm that is native to tensor decompositions and avoids working with the full high-resolution data explicitly. Instead, it actively selects local representative samples that we fetch out-of-core and on demand. The required number of samples grows only logarithmically with the size of the input. Our implicit representation of the tensor in the network enables processing large grids that could not be otherwise tractable in their uncompressed form. The proposed approach is particularly useful for large-scale multidimensional grid data (e.g., 3D tomography), and for tasks that require context over a large receptive field (e.g., predicting the medical condition of entire organs). The code is available at https://github.com/aelphy/c-pic.
Mikhail Usvyatsov, Anastasia Makarova, Rafael Ballester-Ripoll, Maksim Rakhuba, Andreas Krause 0001, Konrad Schindler
ICCV6
2021 PC2WF: 3D Wireframe Reconstruction from Raw Point Clouds
Yujia Liu 0002, Stefano D'Aronco, Konrad Schindler, Jan Dirk Wegner
ICLR3
2021 Walk2Map: Extracting Floor Plans from Indoor Walk Trajectories
abstract
Abstract Recent years have seen a proliferation of new digital products for the efficient management of indoor spaces, with important applications like emergency management, virtual property showcasing and interior design. While highly innovative and effective, these products rely on accurate 3D models of the environments considered, including information on both architectural and non‐permanent elements. These models must be created from measured data such as RGB‐D images or 3D point clouds, whose capture and consolidation involves lengthy data workflows. This strongly limits the rate at which 3D models can be produced, preventing the adoption of many digital services for indoor space management. We provide a radical alternative to such data‐intensive procedures by presentingWalk2Map, a data‐driven approach to generate floor plans only from trajectories of a person walking inside the rooms. Thanks to recent advances in data‐driven inertial odometry, such minimalistic input data can be acquired from the IMU readings of consumer‐level smartphones, which allows for an effortless and scalable mapping of real‐world indoor spaces. Our work is based on learning the latent relation between an indoor walk trajectory and the information represented in a floor plan: interior space footprint, portals, and furniture. We distinguish between recovering area‐related (interior footprint, furniture) and wall‐related (doors) information and use two different neural architectures for the two tasks: an image‐based Encoder‐Decoder and a Graph Convolutional Network, respectively. We train our networks using scanned 3D indoor models and apply them in a cascaded fashion on an indoor walk trajectory at inference time. We perform a qualitative and quantitative evaluation using both trajectories simulated from scanned models of interiors and measured, real‐world trajectories, and compare against a baseline method for image‐to‐image translation. The experiments confirm that our technique is viable and allows recovering reliable floor plans from minimal walk trajectory data.
Claudio Mura, Renato Pajarola, Konrad Schindler, Niloy J. Mitra
Comput. Graph. Forum3
2021 MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking
abstract
Abstract Standardized benchmarks have been crucial in pushing the performance of computer vision algorithms, especially since the advent of deep learning. Although leaderboards should not be over-claimed, they often provide the most objective measure of performance and are therefore important guides for research. We present MOTChallenge , a benchmark for single-camera Multiple Object Tracking (MOT) launched in late 2014, to collect existing and new data and create a framework for the standardized evaluation of multiple object tracking methods. The benchmark is focused on multiple people tracking, since pedestrians are by far the most studied object in the tracking community, with applications ranging from robot navigation to self-driving cars. This paper collects the first three releases of the benchmark: (i) MOT15 , along with numerous state-of-the-art results that were submitted in the last years, (ii) MOT16 , which contains new challenging videos, and (iii) MOT17 , that extends MOT16 sequences with more precise labels and evaluates tracking performance on three different object detectors. The second and third release not only offers a significant increase in the number of labeled boxes, but also provide labels for multiple object classes beside pedestrians, as well as the level of visibility for every single object of interest. We finally provide a categorization of state-of-the-art trackers and a broad error analysis. This will help newcomers understand the related work and research trends in the MOT community, and hopefully shed some light into potential future research directions.
Patrick Dendorfer, Aljosa Osep, Anton Milan, Konrad Schindler, Daniel Cremers, Ian D. Reid 0001, Stefan Roth 0001, Laura Leal-Taixé
Int. J. Comput. Vis.4
2021 Guest Editorial: Special Issue on Performance Evaluation in Computer Vision
Daniel Scharstein, Angela Dai, Daniel Kondermann, Torsten Sattler, Konrad Schindler
Int. J. Comput. Vis.5
2020 KAPLAN: A 3D Point Descriptor for Shape Completion
abstract
We present a novel 3D shape completion method that operates directly on unstructured point clouds, thus avoiding resource-intensive data structures like voxel grids. To this end, we introduce KAPLAN, a 3D point descriptor that aggregates local shape information via a series of 2D convolutions. The key idea is to project the points in a local neighborhood onto multiple planes with different orientations. In each of those planes, point properties like normals or point-to-plane distances are aggregated into a 2D grid and abstracted into a feature representation with an efficient 2D convolutional encoder. Since all planes are encoded jointly, the resulting representation nevertheless can capture their correlations and retains knowledge about the underlying 3D shape, without expensive 3D convolutions. Experiments on public datasets show that KAPLAN achieves state-of-the-art performance for 3D shape completion.
Audrey Richard, Ian Cherabier, Martin R. Oswald, Marc Pollefeys, Konrad Schindler
3DV5
2020 Chained Representation Cycling: Learning to Estimate 3D Human Pose and Shape by Cycling Between Representations
abstract
The goal of many computer vision systems is to transform image pixels into 3D representations. Recent popular models use neural networks to regress directly from pixels to 3D object parameters. Such an approach works well when supervision is available, but in problems like human pose and shape estimation, it is difficult to obtain natural images with 3D ground truth. To go one step further, we propose a new architecture that facilitates unsupervised, or lightly supervised, learning. The idea is to break the problem into a series of transformations between increasingly abstract representations. Each step involves a cycle designed to be learnable without annotated training data, and the chain of cycles delivers the final solution. Specifically, we use 2D body part segments as an intermediate representation that contains enough information to be lifted to 3D, and at the same time is simple enough to be learned in an unsupervised way. We demonstrate the method by learning 3D human pose and shape from un-paired and un-annotated images. We also explore varying amounts of paired data and show that cycling greatly alleviates the need for paired data. While we present results for modeling humans, our formulation is general and can be applied to other vision problems.
Nadine Bertsch, Christoph Lassner, Michael J. Black, Konrad Schindler
AAAI4
2020 From Two Rolling Shutters to One Global Shutter
abstract
Most consumer cameras are equipped with electronic rolling shutter, leading to image distortions when the camera moves during image capture. We explore a surprisingly simple camera configuration that makes it possible to undo the rolling shutter distortion: two cameras mounted to have different rolling shutter directions. Such a setup is easy and cheap to build and it possesses the geometric constraints needed to correct rolling shutter distortion using only a sparse set of point correspondences between the two images. We derive equations that describe the underlying geometry for general and special motions and present an efficient method for finding their solutions. Our synthetic and real experiments demonstrate that our approach is able to remove large rolling shutter distortions of all types without relying on any specific scene structure.
Cenek Albl, Zuzana Kukelova, Viktor Larsson, Michal Polic, Tomás Pajdla, Konrad Schindler
CVPR6
2020 Minimal Rolling Shutter Absolute Pose with Unknown Focal Length and Radial Distortion
Zuzana Kukelova, Cenek Albl, Akihiro Sugimoto, Konrad Schindler, Tomás Pajdla
ECCV (5)4
2020 Indoor Scene Recognition in 3D
abstract
Recognising in what type of environment one is located is an important perception task. For instance, for a robot operating indoors it is helpful to be aware whether it is in a kitchen, a hallway or a bedroom. Existing approaches attempt to classify the scene based on 2D images or 2.5D range images. Here, we study scene recognition from 3D point cloud (or voxel) data, and show that it greatly outperforms methods based on 2D birds-eye views. Moreover, we advocate multi-task learning as a way to improve scene recognition, building on the fact that the scene type is highly correlated with the objects in the scene, and therefore with its semantic segmentation into different object classes. In a series of ablation studies, we show that successful scene recognition is not just the recognition of individual objects unique to some scene type (such as a bathtub), but depends on several different cues, including coarse 3D geometry, colour, and the (implicit) distribution of object categories. Moreover, we demonstrate that surprisingly sparse 3D data is sufficient to classify indoor scenes with good accuracy.
Mikhail Usvyatsov, Konrad Schindler
IROS3
2020 Reconstruction of 3D ight trajectories from ad-hoc camera networks
abstract
We present a method to reconstruct the 3D trajectory of an airborne robotic system only from videos recorded with cameras that are unsynchronized, may feature rolling shutter distortion, and whose viewpoints are unknown. Our approach enables robust and accurate outside-in tracking of dynamically flying targets, with cheap and easy-to-deploy equipment. We show that, in spite of the weakly constrained setting, recent developments in computer vision make it possible to reconstruct trajectories in 3D from unsynchronized, uncalibrated networks of consumer cameras, and validate the proposed method in a realistic field experiment. We make our code available along with the data, including cm-accurate groundtruth from differential GNSS navigation.
Jingtong Li, Jesse Murray, Dorina Ismaili, Konrad Schindler, Cenek Albl
IROS4
2020 Inference, Learning and Attention Mechanisms that Exploit and Preserve Sparsity in CNNs
Timo Hackel, Mikhail Usvyatsov, Silvano Galliani, Jan Dirk Wegner, Konrad Schindler
Int. J. Comput. Vis.5
2020 Guest Editorial: Special Issue on ACCV 2018
C. V. Jawahar, Hongdong Li, Greg Mori, Konrad Schindler
Int. J. Comput. Vis.4
2020 3D Fluid Flow Estimation with Integrated Particle Reconstruction
Katrin Lasinger, Christoph Vogel, Thomas Pock, Konrad Schindler
Int. J. Comput. Vis.4
2019 Learned Multi-View Texture Super-Resolution
abstract
We present a super-resolution method capable of creating a high-resolution texture map for a virtual 3D object from a set of lower-resolution images of that object. Our architecture unifies the concepts of (i) multi-view super-resolution based on the redundancy of overlapping views and (ii) single-view super-resolution based on a learned prior of high-resolution (HR) image structure. The principle of multi-view super-resolution is to invert the image formation process and recover the latent HR texture from multiple lower-resolution projections. We map that inverse problem into a block of suitably designed neural network layers, and combine it with a standard encoder-decoder network for learned single-image super-resolution. Wiring the image formation model into the network avoids having to learn perspective mapping from textures to images, and elegantly handles a varying number of input views. Experiments demonstrate that the combination of multi-view observations and learned prior yields improved texture maps.
Audrey Richard, Ian Cherabier, Martin R. Oswald, Vagia Tsiminaki, Marc Pollefeys, Konrad Schindler
3DV6
2019 Guided Super-Resolution As Pixel-to-Pixel Transformation
abstract
Guided super-resolution is a unifying framework for several computer vision tasks where the inputs are a low-resolution source image of some target quantity (e.g., perspective depth acquired with a time-of-flight camera) and a high-resolution guide image from a different domain (e.g., a grey-scale image from a conventional camera); and the target output is a high-resolution version of the source (in our example, a high-res depth map). The standard way of looking at this problem is to formulate it as a super-resolution task, i.e., the source image is upsampled to the target resolution, while transferring the missing high-frequency details from the guide. Here, we propose to turn that interpretation on its head and instead see it as a pixel-to-pixel mapping of the guide image to the domain of the source image. The pixel-wise mapping is parametrised as a multi-layer perceptron, whose weights are learned by minimising the discrepancies between the source image and the downsampled target image. Importantly, our formulation makes it possible to regularise only the mapping function, while avoiding regularisation of the outputs; thus producing crisp, natural-looking images. The proposed method is unsupervised, using only the specific source and guide images to fit the mapping. We evaluate our method on two different tasks, super-resolution of depth maps and of tree height maps. In both cases, we clearly outperform recent baselines in quantitative comparisons, while delivering visually much sharper outputs.
Riccardo de Lutio, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler
ICCV4
2019 Visual recognition in the wild by sampling deep similarity functions
abstract
Recognising relevant objects or object states in its environment is a basic capability for an autonomous robot. The dominant approach to object recognition in images and range images is classification by supervised machine learning, nowadays mostly with deep convolutional neural networks (CNNs). This works well for target classes whose variability can be completely covered with training examples. However, a robot moving in the wild, i.e., in an environment that is not known at the time the recognition system is trained, will often face domain shift: the training data cannot be assumed to exhaustively cover all the within-class variability that will be encountered in the test data. In that situation, learning is in principle possible, since the training set does capture the defining properties, respectively dissimilarities, of the target classes. But directly training a CNN to predict class probabilities is prone to overfitting to irrelevant correlations between the class labels and the specific subset of the target class that is represented in the training set. We explore the idea to instead learn a Siamese CNN that acts as similarity function between pairs of training examples. Class predictions are then obtained by measuring the similarities between a new test instance and the training samples. We show that the CNN embedding correctly recovers the relative similarities to arbitrary class exemplars in the training set. And that therefore few, randomly picked training exemplars are sufficient to achieve good predictions, making the procedure efficient.
Mikhail Usvyatsov, Konrad Schindler
ICRA2
2017 Online Multi-Target Tracking Using Recurrent Neural Networks
abstract
We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, including a) an a-priori unknown and time-varying number of targets, b) a continuous state estimation of all present targets, and c) a discrete combinatorial problem of data association. Most previous methods involve complex models that require tedious tuning of parameters. Here, we propose for the first time, an end-to-end learning approach for online multi-target tracking. Existing deep learning methods are not designed for the above challenges and cannot be trivially applied to the task. Our solution addresses all of the above points in a principled way. Experiments on both synthetic and real data show promising results obtained at ~300 Hz on a standard CPU, and pave the way towards future research in this direction.
Anton Milan, Seyed Hamid Rezatofighi, Anthony R. Dick, Ian D. Reid 0001, Konrad Schindler
AAAI5
2017 Semantic 3D Reconstruction with Finite Element Bases
Audrey Richard, Christoph Vogel, Maros Blaha, Thomas Pock, Konrad Schindler
BMVC5
2017 A Multi-view Stereo Benchmark with High-Resolution Images and Multi-camera Videos
abstract
Motivated by the limitations of existing multi-view stereo benchmarks, we present a novel dataset for this task. Towards this goal, we recorded a variety of indoor and outdoor scenes using a high-precision laser scanner and captured both high-resolution DSLR imagery as well as synchronized low-resolution stereo videos with varying fields-of-view. To align the images with the laser scans, we propose a robust technique which minimizes photometric errors conditioned on the geometry. In contrast to previous datasets, our benchmark provides novel challenges and covers a diverse set of viewpoints and scene types, ranging from natural scenes to man-made indoor and outdoor environments. Furthermore, we provide data at significantly higher temporal and spatial resolution. Our benchmark is the first to cover the important use case of hand-held mobile devices while also providing high-resolution DSLR camera images. We make our datasets and an online evaluation server available at http://www.eth3d.net.
Thomas Schöps, Johannes L. Schönberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, Andreas Geiger 0001
CVPR5
2017 Semantically Informed Multiview Surface Refinement
abstract
We present a method to jointly refine the geometry and semantic segmentation of 3D surface meshes. Our method alternates between updating the shape and the semantic labels. In the geometry refinement step, the mesh is deformed with variational energy minimization, such that it simultaneously maximizes photo-consistency and the compatibility of the semantic segmentations across a set of calibrated images. Label-specific shape priors account for interactions between the geometry and the semantic labels in 3D. In the semantic segmentation step, the labels on the mesh are updated with MRF inference, such that they are compatible with the semantic segmentations in the input images. Also, this step includes prior assumptions about the surface shape of different semantic classes. The priors induce a tight coupling, where semantic information influences the shape update and vice versa. Specifically, we introduce priors that favor (i) adaptive smoothing, depending on the class label; (ii) straightness of class boundaries; and (iii) semantic labels that are consistent with the surface orientation. The novel mesh-based reconstruction is evaluated in a series of experiments with real and synthetic data. We compare both to state-of-the-art, voxel-based semantic 3D reconstruction, and to purely geometric mesh refinement, and demonstrate that the proposed scheme yields improved 3D geometry as well as an improved semantic segmentation.
Maros Blaha, Mathias Rothermel, Martin R. Oswald, Torsten Sattler, Audrey Richard, Jan Dirk Wegner, Marc Pollefeys, Konrad Schindler
ICCV8
2017 Learned Multi-patch Similarity
abstract
Estimating a depth map from multiple views of a scene is a fundamental task in computer vision. As soon as more than two viewpoints are available, one faces the very basic question how to measure similarity across >2 image patches. Surprisingly, no direct solution exists, instead it is common to fall back to more or less robust averaging of two-view similarities. Encouraged by the success of machine learning, and in particular convolutional neural networks, we propose to learn a matching function which directly maps multiple image patches to a scalar similarity score. Experiments on several multi-view datasets demonstrate that this approach has advantages over methods based on pairwise patch similarity.
Wilfried Hartmann, Silvano Galliani, Michal Havlena, Luc Van Gool, Konrad Schindler
ICCV5
2017 Volumetric Flow Estimation for Incompressible Fluids Using the Stationary Stokes Equations
abstract
In experimental fluid dynamics, the flow in a volume of fluid is observed by injecting high-contrast tracer particles and tracking them in multi-view video. Fluid dynamics researchers have developed variants of space-carving to reconstruct the 3D particle distribution at a given time-step, and then use relatively simple local matching to recover the motion over time. On the contrary, estimating the optical flow between two consecutive images is a long-standing standard problem in computer vision, but only little work exists about volumetric 3D flow. Here, we propose a variational method for 3D fluid flow estimation from multi-view data. We start from a 3D version of the standard variational flow model, and investigate different regularization schemes that ensure divergence-free flow fields, to account for the physics of incompressible fluids. Moreover, we propose a semi-dense formulation, to cope with the computational demands of large volumetric datasets. Flow is estimated and regularized at a lower spatial resolution, while the data term is evaluated at full resolution to preserve the discriminative power and geometric precision of the local particle distribution. Extensive experiments reveal that a simple sum of squared differences (SSD) is the most suitable data term for our application. For regularization, an energy whose Euler-Lagrange equations correspond to the stationary Stokes equations leads to the best results. This strictly enforces a divergence-free flow and additionally penalizes the squared gradient of the flow.
Katrin Lasinger, Christoph Vogel, Konrad Schindler
ICCV3
2017 Semantic segmentation of aerial images with explicit class-boundary modeling
abstract
In this work we propose an end-to-end trainable supervised Deep Convolutional Neural Network (DCNN) targeting the task of semantic-segmentation with the addition of class-aware boundary detection. Through this explicit modeling of the class-boundaries, we enforce the network to extract coherent and complete objects, suppressing the uncertainty influencing these regions. Importantly, we show that class-boundary networks in conjunction with DCNN performs optimally, achieving over 90% overall accuracy (OA) on the challenging ISPRS Vaihingen Semantic Segmentation benchmark.
Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner, Mihai Datcu, Uwe Stilla
IGARSS2
2017 Gaze-Informed location-based services
abstract
Location-based services (LBS) provide more useful, intelligent assistance to users by adapting to their geographic context. For some services that context goes beyond a location and includes further spatial parameters, such as the user’s orientation or field of view. Here, we introduce Gaze-Informed LBS (GAIN-LBS), a novel type of LBS that takes into account the user’s viewing direction. Such a system could, for instance, provide audio information about the specific building a tourist is looking at from a vantage point. To determine the viewing direction relative to the environment, we record the gaze direction relative to the user’s head with a mobile eye tracker. Image data from the tracker’s forward-looking camera serve as input to determine the orientation of the head w.r.t. the surrounding scene, using computer vision methods that allow one to estimate the relative transformation between the camera and a known view of the scene in real-time and without the need for artificial markers or additional sensors. We focus on how to map the point of regard of a user to a reference system, for which the objects of interest are known in advance. In an experimental validation on three real city panoramas, we confirm that the approach can cope with head movements of varying speed, including fast rotations up to to 63 degrees per second. We further demonstrate the feasibility of GAIN-LBS for tourist assistance with a proof-of-concept experiment in which a tourist explores a city panorama, where the approach achieved a recall that reaches over 99%. Finally, a GAIN-LBS can provide objective and qualitative ways of examining the gaze of a user based on what the user is currently looking at.
Vasileios Athanasios Anagnostopoulos, Michal Havlena, Peter Kiefer, Ioannis Giannopoulos, Konrad Schindler, Martin Raubal
Int. J. Geogr. Inf. Sci.5
2017 Learning Aerial Image Segmentation From Online Maps
abstract
This paper deals with semantic segmentation of high-resolution (aerial) images where a semantic class label is assigned to each pixel via supervised classification as a basis for automatic map generation. Recently, deep convolutional neural networks (CNNs) have shown impressive performance and have quickly become the de-facto standard for semantic segmentation, with the added benefit that task-specific feature design is no longer necessary. However, a major downside of deep learning methods is that they are extremely data hungry, thus aggravating the perennial bottleneck of supervised classification, to obtain enough annotated training data. On the other hand, it has been observed that they are rather robust against noise in the training labels. This opens up the intriguing possibility to avoid annotating huge amounts of training data, and instead train the classifier from existing legacy data or crowd-sourced maps that can exhibit high levels of noise. The question addressed in this paper is: can training with large-scale publicly available labels replace a substantial part of the manual labeling effort and still achieve sufficient performance? Such data will inevitably contain a significant portion of errors, but in return virtually unlimited quantities of it are available in larger parts of the world. We adapt a state-of-the-art CNN architecture for semantic segmentation of buildings and roads in aerial images, and compare its performance when using different training data sets, ranging from manually labeled pixel-accurate ground truth of the same city to automatic training data derived from OpenStreetMap data from distant locations. We report our results that indicate that satisfying performance can be obtained with significantly less manual annotation effort, by exploiting noisy large-scale training data.
Pascal Kaiser, Jan Dirk Wegner, Aurélien Lucchi, Martin Jaggi, Thomas Hofmann 0001, Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.6
2016 Large-Scale Semantic 3D Reconstruction: An Adaptive Multi-resolution Model for Multi-class Volumetric Labeling
abstract
We propose an adaptive multi-resolution formulation of semantic 3D reconstruction. Given a set of images of a scene, semantic 3D reconstruction aims to densely reconstruct both the 3D shape of the scene and a segmentation into semantic object classes. Jointly reasoning about shape and class allows one to take into account class-specific shape priors (e.g., building walls should be smooth and vertical, and vice versa smooth, vertical surfaces are likely to be building walls), leading to improved reconstruction results. So far, semantic 3D reconstruction methods have been limited to small scenes and low resolution, because of their large memory footprint and computational cost. To scale them up to large scenes, we propose a hierarchical scheme which refines the reconstruction only in regions that are likely to contain a surface, exploiting the fact that both high spatial resolution and high numerical precision are only required in those regions. Our scheme amounts to solving a sequence of convex optimizations while progressively removing constraints, in such a way that the energy, in each iteration, is the tightest possible approximation of the underlying energy at full resolution. In our experiments the method saves up to 98% memory and 95% computation time, without any loss of accuracy.
Maros Blaha, Christoph Vogel, Audrey Richard, Jan Dirk Wegner, Thomas Pock, Konrad Schindler
CVPR6
2016 Just Look at the Image: Viewpoint-Specific Surface Normal Prediction for Improved Multi-View Reconstruction
abstract
We present a multi-view reconstruction method that combines conventional multi-view stereo (MVS) with appearance-based normal prediction, to obtain dense and accurate 3D surface models. Reliable surface normals reconstructed from multi-view correspondence serve as training data for a convolutional neural network (CNN), which predicts continuous normal vectors from raw image patches. By training from known points in the same image, the prediction is specifically tailored to the materials and lighting conditions of the particular scene, as well as to the precise camera viewpoint. It is therefore a lot easier to learn than generic single-view normal estimation. The estimated normal maps, together with the known depth values from MVS, are integrated to dense depth maps, which in turn are fused into a 3D model. Experiments on the DTU dataset show that our method delivers 3D reconstructions with the same accuracy as MVS, but with significantly higher completeness.
Silvano Galliani, Konrad Schindler
CVPR2
2016 Contour Detection in Unstructured 3D Point Clouds
abstract
We describe a method to automatically detect contours, i.e. lines along which the surface orientation sharply changes, in large-scale outdoor point clouds. Contours are important intermediate features for structuring point clouds and converting them into high-quality surface or solid models, and are extensively used in graphics and mapping applications. Yet, detecting them in unstructured, inhomogeneous point clouds turns out to be surprisingly difficult, and existing line detection algorithms largely fail. We approach contour extraction as a two-stage discriminative learning problem. In the first stage, a contour score for each individual point is predicted with a binary classifier, using a set of features extracted from the point's neighborhood. The contour scores serve as a basis to construct an overcomplete graph of candidate contours. The second stage selects an optimal set of contours from the candidates. This amounts to a further binary classification in a higher-order MRF, whose cliques encode a preference for connected contours and penalize loose ends. The method can handle point clouds > 107 points in a couple of minutes, and vastly outperforms a baseline that performs Canny-style edge detection on a range image representation of the point cloud.
Timo Hackel, Jan Dirk Wegner, Konrad Schindler
CVPR3
2016 Large-Scale Location Recognition and the Geometric Burstiness Problem
abstract
Visual location recognition is the task of determining the place depicted in a query image from a given database of geo-tagged images. Location recognition is often cast as an image retrieval problem and recent research has almost exclusively focused on improving the chance that a relevant database image is ranked high enough after retrieval. The implicit assumption is that the number of inliers found by spatial verification can be used to distinguish between a related and an unrelated database photo with high precision. In this paper, we show that this assumption does not hold for large datasets due to the appearance of geometric bursts, i.e., sets of visual elements appearing in similar geometric configurations in unrelated database photos. We propose algorithms for detecting and handling geometric bursts. Although conceptually simple, using the proposed weighting schemes dramatically improves the recall that can be achieved when high precision is required compared to the standard re-ranking based on the inlier count. Our approach is easy to implement and can easily be integrated into existing location recognition systems.
Torsten Sattler, Michal Havlena, Konrad Schindler, Marc Pollefeys
CVPR3
2016 Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees
abstract
Each corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design computer vision systems that will help us search, catalog and monitor public infrastructure, buildings and artifacts. We explore the architecture and feasibility of such a system. The main technical challenge is combining test time information from multiple views of each geographic location (e.g., aerial and street views). We implement two modules: det2geo, which detects the set of locations of objects belonging to a given category, and geo2cat, which computes the fine-grained category of the object at a given location. We introduce a solution that adapts state-of the-art CNN-based object detectors and classifiers. We test our method on "Pasadena Urban Trees", a new dataset of 80,000 trees with geographic and species annotations, and show that combining multiple views significantly improves both tree detection and tree species classification, rivaling human performance.
Jan Dirk Wegner, Steve Branson, David Hall 0002, Konrad Schindler, Pietro Perona
CVPR4
2016 Multi-Target Tracking by Discrete-Continuous Energy Minimization
abstract
The task of tracking multiple targets is often addressed with the so-called tracking-by-detection paradigm, where the first step is to obtain a set of target hypotheses for each frame independently. Tracking can then be regarded as solving two separate, but tightly coupled problems. The first is to carry out data association, i.e., to determine the origin of each of the available observations. The second problem is to reconstruct the actual trajectories that describe the spatio-temporal motion pattern of each individual target. The former is inherently a discrete problem, while the latter should intuitively be modeled in continuous space. Having to deal with an unknown number of targets, complex dependencies, and physical constraints, both are challenging tasks on their own and thus most previous work focuses on one of these subproblems. Here, we present a multi-target tracking approach that explicitly models both tasks as minimization of a unified discrete-continuous energy function. Trajectory properties are captured through global label costs, a recent concept from multi-model fitting, which we introduce to tracking. Specifically, label costs describe physical properties of individual tracks, e.g., linear and angular dynamics, or entry and exit points. We further introduce pairwise label costs to describe mutual interactions between targets in order to avoid collisions. By choosing appropriate forms for the individual energy components, powerful discrete optimization techniques can be leveraged to address data association, while the shapes of individual trajectories are updated by gradient-based continuous energy minimization. The proposed method achieves state-of-the-art results on diverse benchmark sequences.
Anton Milan, Konrad Schindler, Stefan Roth 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Joint tracking and segmentation of multiple targets
abstract
Tracking-by-detection has proven to be the most successful strategy to address the task of tracking multiple targets in unconstrained scenarios [e.g. 40, 53, 55]. Traditionally, a set of sparse detections, generated in a preprocessing step, serves as input to a high-level tracker whose goal is to correctly associate these “dots” over time. An obvious short-coming of this approach is that most information available in image sequences is simply ignored by thresholding weak detection responses and applying non-maximum suppression. We propose a multi-target tracker that exploits low level image information and associates every (super)-pixel to a specific target or classifies it as background. As a result, we obtain a video segmentation in addition to the classical bounding-box representation in unconstrained, real-world videos. Our method shows encouraging results on many standard benchmark sequences and significantly outperforms state-of-the-art tracking-by-detection approaches in crowded scenes with long-term partial occlusions.
Anton Milan, Laura Leal-Taixé, Konrad Schindler, Ian D. Reid 0001
CVPR3
2015 Massively Parallel Multiview Stereopsis by Surface Normal Diffusion
abstract
We present a new, massively parallel method for high-quality multiview matching. Our work builds on the Patchmatch idea: starting from randomly generated 3D planes in scene space, the best-fitting planes are iteratively propagated and refined to obtain a 3D depth and normal field per view, such that a robust photo-consistency measure over all images is maximized. Our main novelties are on the one hand to formulate Patchmatch in scene space, which makes it possible to aggregate image similarity across multiple views and obtain more accurate depth maps. And on the other hand a modified, diffusion-like propagation scheme that can be massively parallelized and delivers dense multiview correspondence over ten 1.9-Megapixel images in 3 seconds, on a consumer-grade GPU. Our method uses a slanted support window and thus has no fronto-parallel bias, it is completely local and parallel, such that computation time scales linearly with image size, and inversely proportional to the number of parallel threads. Furthermore, it has low memory footprint (four values per pixel, independent of the depth range). It therefore scales exceptionally well and can handle multiple large images at high depth resolution. Experiments on the DTU and Middlebury multiview datasets as well as oblique aerial images show that our method achieves very competitive results with high accuracy and completeness, across a range of different scenarios.
Silvano Galliani, Katrin Lasinger, Konrad Schindler
ICCV3
2015 Hyperspectral Super-Resolution by Coupled Spectral Unmixing
abstract
Hyperspectral cameras capture images with many narrow spectral channels, which densely sample the electromagnetic spectrum. The detailed spectral resolution is useful for many image analysis problems, but it comes at the cost of much lower spatial resolution. Hyperspectral super-resolution addresses this problem, by fusing a low-resolution hyperspectral image and a conventional high-resolution image into a product of both high spatial and high spectral resolution. In this paper, we propose a method which performs hyperspectral super-resolution by jointly unmixing the two input images into the pure reflectance spectra of the observed materials and the associated mixing coefficients. The formulation leads to a coupled matrix factorisation problem, with a number of useful constraints imposed by elementary physical properties of spectral mixing. In experiments with two benchmark datasets we show that the proposed approach delivers improved hyperspectral super-resolution.
Charis Lanaras, Emmanuel Baltsavias, Konrad Schindler
ICCV3
2015 Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition
abstract
Structure-based localization is the task of finding the absolute pose of a given query image w.r.t. a pre-computed 3D model. While this is almost trivial at small scale, special care must be taken as the size of the 3D model grows, because straight-forward descriptor matching becomes ineffective due to the large memory footprint of the model, as well as the strictness of the ratio test in 3D. Recently, several authors have tried to overcome these problems, either by a smart compression of the 3D model or by clever sampling strategies for geometric verification. Here we explore an orthogonal strategy, which uses all the 3D points and standard sampling, but performs feature matching implicitly, by quantization into a fine vocabulary. We show that although this matching is ambiguous and gives rise to 3D hyperpoints when matching each 2D query feature in isolation, a simple voting strategy, which enforces the fact that the selected 3D points shall be co-visible, can reliably find a locally unique 2D-3D point assignment. Experiments on two large-scale datasets demonstrate that our method achieves state-of-the-art performance, while the memory footprint is greatly reduced, since only visual word labels but no 3D point descriptors need to be stored.
Torsten Sattler, Michal Havlena, Filip Radenovic, Konrad Schindler, Marc Pollefeys
ICCV4
2015 Pose Estimation of Object Categories in Videos Using Linear Programming
abstract
In this paper we propose a method to consistently recover the pose of an object from a known class in a video sequence. As individual poses estimated from monocular images are rather noisy, we optimally aggregate pose evidence over all video frames. We construct a graph where nodes are values sampled from the pose posterior distributions computed by a continuous pose estimator in each frame of the sequence. We then find the globally optimum pose path through the graph that best explains the pose evidence for the whole sequence. As a result, we recover the correct object orientation at each frame even if single-frame pose evidence is sometimes inaccurate. We evaluate our approach on two publicly available car datasets, which encompass busy street scenarios and car races with significant changes in car orientation, blur and occlusions. We show that our method outperforms state-of-the-art approaches reducing the error by 40% on the challenging KITTI dataset.
Michele Fenzi, Laura Leal-Taixé, Konrad Schindler, Jörn Ostermann
WACV3
2015 Visual Gyroscope for Accurate Orientation Estimation
abstract
A visual gyroscope is a device which estimates camera 3D rotation using image input, in our case a monocular video. Contrary to traditional Structure-From-Motion (SFM) or visual SLAM, we address the case where only the rotation must be found, whereas no translation estimate is desired. That case can be solved without computing an explicit 3D map of the environment, thus avoiding computationally expensive bundle adjustment. Instead, a simple linear method is used to obtain globally consistent rotations from relative rotation estimates between image pairs. We show that the obtained camera orientations are accurate w.r.t. ground truth collected with a navigation-grade (dGPS-supported) IMU, and reach 3D bearing errors below 1 over a 1-minute time interval for >90% of all cases. Efficient computation is achieved by employing GPU-enabled feature extraction and matching. To warrant on-line performance for sequences of arbitrary lengths we run the global rotation estimation in a sliding-window fashion and show that the accuracy of the camera orientations obtained by chaining the partial solutions stays high. Finally, we compare the proposed visual gyroscope to a publicly available SFM software and experimentally demonstrate the importance of a very large field-of-view for accurate rotation estimation.
Wilfried Hartmann, Michal Havlena, Konrad Schindler
WACV3
2015 3D Scene Flow Estimation with a Piecewise Rigid Scene Model
Christoph Vogel, Konrad Schindler, Stefan Roth 0001
Int. J. Comput. Vis.2
2015 Towards Scene Understanding with Detailed 3D Object Representations
M. Zeeshan Zia, Michael Stark 0003, Konrad Schindler
Int. J. Comput. Vis.3
2015 Features, Color Spaces, and Boosting: New Insights on Semantic Classification of Remote Sensing Images
abstract
A major yet largely unsolved problem in the semantic classification of very high resolution remote sensing images is the design and selection of appropriate features. At a ground sampling distance below half a meter, fine-grained texture details of objects emerge and lead to a large intraclass variability while generally keeping the between-class variability at a low level. Usually, the user makes an educated guess on what features seem to appropriately capture characteristic object class patterns. Here, we propose to avoid manual feature selection and let a boosting classifier choose optimal features from a vast Randomized Quasi-Exhaustive (RQE) set of feature candidates directly during training. This RQE feature set consists of a multitude of very simple features that are computed efficiently via integral images inside a sliding window. This simple but comprehensive feature candidate set enables the boosting classifier to assemble the most discriminative textures at different scale levels to classify a small number of broad urban land-cover classes. We do an extensive evaluation on several data sets and compare performance against multiple feature extraction baselines in different color spaces. In addition, we verify experimentally if we gain any classification accuracy if moving from boosting stumps to trees. Cross-validation minimizes the possible bias caused by specific training/testing setups. It turns out that boosting in combination with the proposed RQE feature set outperforms all baseline features while still remaining computationally efficient. Particularly boosting trees (instead of stumps) captures class patterns so well that results suggest to completely leave feature selection to the classifier.
Piotr Tokarczyk, Jan Dirk Wegner, Stefan Walk, Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.4
2014 Predicting Matchability
abstract
The initial steps of many computer vision algorithms are interest point extraction and matching. In larger image sets the pairwise matching of interest point descriptors between images is an important bottleneck. For each descriptor in one image the (approximate) nearest neighbor in the other one has to be found and checked against the second-nearest neighbor to ensure the correspondence is unambiguous. Here, we asked the question how to best decimate the list of interest points without losing matches, i.e. we aim to speed up matching by filtering out, in advance, those points which would not survive the matching stage. It turns out that the best filtering criterion is not the response of the interest point detector, which in fact is not surprising: the goal of detection are repeatable and well-localized points, whereas the objective of the selection are points whose descriptors can be matched successfully. We show that one can in fact learn to predict which descriptors are matchable, and thus reduce the number of interest points significantly without losing too many matches. We show that this strategy, as simple as it is, greatly improves the matching success with the same number of points per image. Moreover, we embed the prediction in a state-of-the-art Structure-from-Motion pipeline and demonstrate that it also outperforms other selection methods at system level.
Wilfried Hartmann, Michal Havlena, Konrad Schindler
CVPR3
2014 Are Cars Just 3D Boxes? Jointly Estimating the 3D Shape of Multiple Objects
abstract
Current systems for scene understanding typically represent objects as 2D or 3D bounding boxes. While these representations have proven robust in a variety of applications, they provide only coarse approximations to the true 2D and 3D extent of objects. As a result, object-object interactions, such as occlusions or ground-plane contact, can be represented only superficially. In this paper, we approach the problem of scene understanding from the perspective of 3D shape modeling, and design a 3D scene representation that reasons jointly about the 3D shape of multiple objects. This representation allows to express 3D geometry and occlusion on the fine detail level of individual vertices of 3D wireframe models, and makes it possible to treat dependencies between objects, such as occlusion reasoning, in a deterministic way. In our experiments, we demonstrate the benefit of jointly estimating the 3D shape of multiple objects in a scene over working with coarse boxes, on the recently proposed KITTI dataset of realistic street scenes.
M. Zeeshan Zia, Michael Stark 0003, Konrad Schindler
CVPR3
2014 VocMatch: Efficient Multiview Correspondence for Structure from Motion
Michal Havlena, Konrad Schindler
ECCV (3)2
2014 View-Consistent 3D Scene Flow Estimation over Multiple Frames
Christoph Vogel, Stefan Roth 0001, Konrad Schindler
ECCV (4)3
2014 Continuous Energy Minimization for Multitarget Tracking
abstract
Many recent advances in multiple target tracking aim at finding a (nearly) optimal set of trajectories within a temporal window. To handle the large space of possible trajectory hypotheses, it is typically reduced to a finite set by some form of data-driven or regular discretization. In this work, we propose an alternative formulation of multitarget tracking as minimization of a continuous energy. Contrary to recent approaches, we focus on designing an energy that corresponds to a more complete representation of the problem, rather than one that is amenable to global optimization. Besides the image evidence, the energy function takes into account physical constraints, such as target dynamics, mutual exclusion, and track persistence. In addition, partial image evidence is handled with explicit occlusion reasoning, and different targets are disambiguated with an appearance model. To nevertheless find strong local minima of the proposed nonconvex energy, we construct a suitable optimization scheme that alternates between continuous conjugate gradient descent and discrete transdimensional jump moves. These moves, which are executed such that they always reduce the energy, allow the search to escape weak minima and explore a much larger portion of the search space of varying dimensionality. We demonstrate the validity of our approach with an extensive quantitative evaluation on several public data sets.
Anton Milan, Stefan Roth 0001, Konrad Schindler
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Interacting Geometric Priors For Robust Multimodel Fitting
abstract
Recent works on multimodel fitting are often formulated as an energy minimization task, where the energy function includes fitting error and regularization terms, such as low-level spatial smoothness and model complexity. In this paper, we introduce a novel energy with high-level geometric priors that consider interactions between geometric models, such that certain preferred model configurations may be induced.We argue that in many applications, such prior geometric properties are available and should be fruitfully exploited. For example, in surface fitting to point clouds, the building walls are usually either orthogonal or parallel to each other. Our proposed energy function is useful in dealing with unknown distributions of data errors and outliers, which are often the factors leading to biased estimation. Furthermore, the energy can be efficiently minimized using the expansion move method. We evaluate the performance on several vision applications using real data sets. Experimental results show that our method outperforms the state-of-the-art methods without significant increase in computation.
Trung-Thanh Pham, Tat-Jun Chin, Konrad Schindler, David Suter
IEEE Trans. Image Process.3
2013 Detection- and Trajectory-Level Exclusion in Multiple Object Tracking
abstract
When tracking multiple targets in crowded scenarios, modeling mutual exclusion between distinct targets becomes important at two levels: (1) in data association, each target observation should support at most one trajectory and each trajectory should be assigned at most one observation per frame, (2) in trajectory estimation, two trajectories should remain spatially separated at all times to avoid collisions. Yet, existing trackers often sidestep these important constraints. We address this using a mixed discrete-continuous conditional random field (CRF) that explicitly models both types of constraints: Exclusion between conflicting observations with super modular pairwise terms, and exclusion between trajectories by generalizing global label costs to suppress the co-occurrence of incompatible labels (trajectories). We develop an expansion move-based MAP estimation scheme that handles both non-sub modular constraints and pairwise global label costs. Furthermore, we perform a statistical analysis of ground-truth trajectories to derive appropriate CRF potentials for modeling data fidelity, target dynamics, and inter-target occlusion.
Anton Milan, Konrad Schindler, Stefan Roth 0001
CVPR2
2013 A Higher-Order CRF Model for Road Network Extraction
abstract
The aim of this work is to extract the road network from aerial images. What makes the problem challenging is the complex structure of the prior: roads form a connected network of smooth, thin segments which meet at junctions and crossings. This type of a-priori knowledge is more difficult to turn into a tractable model than standard smoothness or co-occurrence assumptions. We develop a novel CRF formulation for road labeling, in which the prior is represented by higher-order cliques that connect sets of super pixels along straight line segments. These long-range cliques have asymmetric PN-potentials, which express a preference to assign all rather than just some of their constituent super pixels to the road class. Thus, the road likelihood is amplified for thin chains of super pixels, while the CRF is still amenable to optimization with graph cuts. Since the number of such cliques of arbitrary length is huge, we furthermore propose a sampling scheme which concentrates on those cliques which are most relevant for the optimization. In experiments on two different databases the model significantly improves both the per-pixel accuracy and the topological correctness of the extracted roads, and outperforms both a simple smoothness prior and heuristic rule-based road completion.
Jan Dirk Wegner, Javier A. Montoya-Zegarra, Konrad Schindler
CVPR3
2013 Explicit Occlusion Modeling for 3D Object Class Representations
abstract
Despite the success of current state-of-the-art object class detectors, severe occlusion remains a major challenge. This is particularly true for more geometrically expressive 3D object class representations. While these representations have attracted renewed interest for precise object pose estimation, the focus has mostly been on rather clean datasets, where occlusion is not an issue. In this paper, we tackle the challenge of modeling occlusion in the context of a 3D geometric object class model that is capable of fine-grained, part-level 3D object reconstruction. Following the intuition that 3D modeling should facilitate occlusion reasoning, we design an explicit representation of likely geometric occlusion patterns. Robustness is achieved by pooling image evidence from of a set of fixed part detectors as well as a non-parametric representation of part configurations in the spirit of pose lets. We confirm the potential of our method on cars in a newly collected data set of inner-city street scenes with varying levels of occlusion, and demonstrate superior performance in occlusion estimation and part localization, compared to baselines that are unaware of occlusions.
M. Zeeshan Zia, Michael Stark 0003, Konrad Schindler
CVPR3
2013 Learning People Detectors for Tracking in Crowded Scenes
abstract
People tracking in crowded real-world scenes is challenging due to frequent and long-term occlusions. Recent tracking methods obtain the image evidence from object (people) detectors, but typically use off-the-shelf detectors and treat them as black box components. In this paper we argue that for best performance one should explicitly train people detectors on failure cases of the overall tracker instead. To that end, we first propose a novel joint people detector that combines a state-of-the-art single person detector with a detector for pairs of people, which explicitly exploits common patterns of person-person occlusions across multiple viewpoints that are a frequent failure case for tracking in crowded scenes. To explicitly address remaining failure modes of the tracker we explore two methods. First, we analyze typical failures of trackers and train a detector explicitly on these cases. And second, we train the detector with the people tracker in the loop, focusing on the most common tracker failures. We show that our joint multi-person detector significantly improves both detection accuracy as well as tracker performance, improving the state-of-the-art on standard benchmarks.
Siyu Tang 0001, Mykhaylo Andriluka, Anton Milan, Konrad Schindler, Stefan Roth 0001, Bernt Schiele
ICCV4
2013 Piecewise Rigid Scene Flow
abstract
Estimating dense 3D scene flow from stereo sequences remains a challenging task, despite much progress in both classical disparity and 2D optical flow estimation. To overcome the limitations of existing techniques, we introduce a novel model that represents the dynamic 3D scene by a collection of planar, rigidly moving, local segments. Scene flow estimation then amounts to jointly estimating the pixel-to-segment assignment, and the 3D position, normal vector, and rigid motion parameters of a plane for each segment. The proposed energy combines an occlusion-sensitive data term with appropriate shape, motion, and segmentation regularizers. Optimization proceeds in two stages: Starting from an initial super pixelization, we estimate the shape and motion parameters of all segments by assigning a proposal from a set of moving planes. Then the pixel-to-segment assignment is updated, while holding the shape and motion parameters of the moving planes fixed. We demonstrate the benefits of our model on different real-world image sets, including the challenging KITTI benchmark. We achieve leading performance levels, exceeding competing 3D scene flow methods, and even yielding better 2D motion estimates than all tested dedicated optical flow techniques.
Christoph Vogel, Konrad Schindler, Stefan Roth 0001
ICCV2
2013 Semantic tie points
abstract
Images for 3D mapping are always recorded in such a way that relevant scene parts are seen from multiple viewpoints, so as to facilitate camera orientation and 3D point triangulation. Beyond geometric reconstruction, automatic mapping also requires the semantic interpretation of the image content, and for that task the redundancy provided by overlapping images has been exploited much less. Here we address the task of learning a classifier for pixel-wise semantic labeling of the observed scene. The main insight is that the mere fact that two regions in different images depict the same 3D scene point yields a constraint which can be exploited in the learning phase, namely that they should receive the same class label, even if it is not known which one. In analogy to geometric “tie points” - image correspondences with a priori unknown 3D coordinates, which nevertheless constrain camera orientation - we call these correspondences “semantic tie points”. We show how to integrate this weaker form of supervision, which is readily available in any multi-view dataset, into a random forest classifier, and demonstrate improved classification performance of the resulting classifier in an aerial dataset.
Javier A. Montoya-Zegarra, Christian Leistner, Konrad Schindler
WACV3
2013 Monocular Visual Scene Understanding: Understanding Multi-Object Traffic Scenes
abstract
Following recent advances in detection, context modeling, and tracking, scene understanding has been the focus of renewed interest in computer vision research. This paper presents a novel probabilistic 3D scene model that integrates state-of-the-art multiclass object detection, object tracking and scene labeling together with geometric 3D reasoning. Our model is able to represent complex object interactions such as inter-object occlusion, physical exclusion between objects, and geometric context. Inference in this model allows us to jointly recover the 3D scene context and perform 3D multi-object tracking from a mobile observer, for objects of multiple categories, using only monocular video as input. Contrary to many other approaches, our system performs explicit occlusion reasoning and is therefore capable of tracking objects that are partially occluded for extended periods of time, or objects that have never been observed to their full extent. In addition, we show that a joint scene tracklet model for the evidence collected over multiple frames substantially improves performance. The approach is evaluated for different types of challenging onboard sequences. We first show a substantial improvement to the state of the art in 3D multipeople tracking. Moreover, a similar performance gain is achieved for multiclass 3D tracking of cars and trucks on a challenging dataset.
Christian Wojek, Stefan Walk, Stefan Roth 0001, Konrad Schindler, Bernt Schiele
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Detailed 3D Representations for Object Recognition and Modeling
abstract
Geometric 3D reasoning at the level of objects has received renewed attention recently in the context of visual scene understanding. The level of geometric detail, however, is typically limited to qualitative representations or coarse boxes. This is linked to the fact that today's object class detectors are tuned toward robust 2D matching rather than accurate 3D geometry, encouraged by bounding-box-based benchmarks such as Pascal VOC. In this paper, we revisit ideas from the early days of computer vision, namely, detailed, 3D geometric object class representations for recognition. These representations can recover geometrically far more accurate object hypotheses than just bounding boxes, including continuous estimates of object pose and 3D wireframes with relative 3D positions of object parts. In combination with robust techniques for shape description and inference, we outperform state-of-the-art results in monocular 3D pose estimation. In a series of experiments, we analyze our approach in detail and demonstrate novel applications enabled by such an object class representation, such as fine-grained categorization of cars and bicycles, according to their 3D geometry, and ultrawide baseline matching.
M. Zeeshan Zia, Michael Stark 0003, Bernt Schiele, Konrad Schindler
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Discrete-continuous optimization for multi-target tracking
abstract
The problem of multi-target tracking is comprised of two distinct, but tightly coupled challenges: (i) the naturally discrete problem of data association, i.e. assigning image observations to the appropriate target; (ii) the naturally continuous problem of trajectory estimation, i.e. recovering the trajectories of all targets. To go beyond simple greedy solutions for data association, recent approaches often perform multi-target tracking using discrete optimization. This has the disadvantage that trajectories need to be pre-computed or represented discretely, thus limiting accuracy. In this paper we instead formulate multi-target tracking as a discrete-continuous optimization problem that handles each aspect in its natural domain and allows leveraging powerful methods for multi-model fitting. Data association is performed using discrete optimization with label costs, yielding near optimality. Trajectory estimation is posed as a continuous fitting problem with a simple closed-form solution, which is used in turn to update the label costs. We demonstrate the accuracy and robustness of our approach with state-of-the-art performance on several standard datasets.
Anton Andriyenko, Konrad Schindler, Stefan Roth 0001
CVPR2
2012 A Computational Feedforward Model Predicts Categorization of Masked Emotional Body Language for Longer, but Not for Shorter, Latencies
abstract
Given the presence of massive feedback loops in brain networks, it is difficult to disentangle the contribution of feedforward and feedback processing to the recognition of visual stimuli, in this case, of emotional body expressions. The aim of the work presented in this letter is to shed light on how well feedforward processing explains rapid categorization of this important class of stimuli. By means of parametric masking, it may be possible to control the contribution of feedback activity in human participants. A close comparison is presented between human recognition performance and the performance of a computational neural model that exclusively modeled feedforward processing and was engineered to fulfill the computational requirements of recognition. Results show that the longer the stimulus onset asynchrony (SOA), the closer the performance of the human participants was to the values predicted by the model, with an optimum at an SOA of 100 ms. At short SOA latencies, human performance deteriorated, but the categorization of the emotional expressions was still above baseline. The data suggest that, although theoretically, feedback arising from inferotemporal cortex is likely to be blocked when the SOA is 100 ms, human participants still seem to rely on more local visual feedback processing to equal the model's performance.
Bernard M. C. Stienen, Konrad Schindler, Béatrice de Gelder
Neural Comput.2
2012 An Overview and Comparison of Smooth Labeling Methods for Land-Cover Classification
abstract
An elementary piece of our prior knowledge about images of the physical world is that they are spatially smooth, in the sense that neighboring pixels are more likely to belong to the same object (class) than to different ones. The smoothness assumption becomes more important as sensor resolutions keep increasing, both because the radiometric variability within classes increases and because remote sensing is employed in more heterogeneous areas (e.g., cities), where shadow and shading effects, a multitude of materials, etc., degrade the measurement data, and prior knowledge plays a greater role. This paper gives a systematic overview of image classification methods, which impose a smoothness prior on the labels. Both local filtering-type approaches and global random field models developed in other fields of image processing are reviewed, and two new methods are proposed. Then follows a detailed experimental comparison and analysis of the presented methods, using two different aerial data sets from urban areas with known ground truth. A main message of the paper is that when classifying data of high spatial resolution, smoothness greatly improves the accuracy of the result-in our experiments up to 33%. A further finding is that global random field models outperform local filtering methods and should be more widely adopted for remote sensing. Finally, the evaluation confirms that all methods already oversmooth when most effective, pointing out that there is a need to include more and more complex prior information into the classification process.
Konrad Schindler
IEEE Trans. Geosci. Remote. Sens.1
2011 Multi-target tracking by continuous energy minimization
abstract
We propose to formulate multi-target tracking as minimization of a continuous energy function. Other than a number of recent approaches we focus on designing an energy function that represents the problem as faithfully as possible, rather than one that is amenable to elegant optimization. We then go on to construct a suitable optimization scheme to find strong local minima of the proposed energy. The scheme extends the conjugate gradient method with periodic trans-dimensional jumps. These moves allow the search to escape weak minima and explore a much larger portion of the variable-dimensional search space, while still always reducing the energy. To demonstrate the validity of this approach we present an extensive quantitative evaluation both on synthetic data and on six different real video sequences. In both cases we achieve a significant performance improvement over an extended Kalman filter baseline as well as an ILP-based state-of-the-art tracker.
Anton Andriyenko, Konrad Schindler
CVPR2
2011 3D scene flow estimation with a rigid motion prior
abstract
We present an approach to 3D scene flow estimation, which exploits that in realistic scenarios image motion is frequently dominated by observer motion and independent, but rigid object motion. We cast the dense estimation of both scene structure and 3D motion from sequences of two or more views as a single energy minimization problem. We show that agnostic smoothness priors, such as the popular total variation, are biased against motion discontinuities in viewing direction. Instead, we propose to regularize by encouraging local rigidity of the 3D scene. We derive a local rigidity constraint of the 3D scene flow and define a smoothness term that penalizes deviations from that constraint, thus favoring solutions that consist largely of rigidly moving parts. Our experiments show that the new rigid motion prior reduces the 3D flow error by 42% compared to standard TV regularization with the same data term.
Christoph Vogel, Konrad Schindler, Stefan Roth 0001
ICCV2
2010 New features and insights for pedestrian detection
abstract
Despite impressive progress in people detection the performance on challenging datasets like Caltech Pedestrians or TUD-Brussels is still unsatisfactory. In this work we show that motion features derived from optic flow yield substantial improvements on image sequences, if implemented correctly - even in the case of low-quality video and consequently degraded flow fields. Furthermore, we introduce a new feature, self-similarity on color channels, which consistently improves detection performance both for static images and for video sequences, across different datasets. In combination with HOG, these two features outperform the state-of-the-art by up to 20%. Finally, we report two insights concerning detector evaluations, which apply to classifier-based object detection in general. First, we show that a commonly under-estimated detail of training, the number of bootstrapping rounds, has a drastic influence on the relative (and absolute) performance of different feature/classifier combinations. Second, we discuss important intricacies of detector evaluation and show that current benchmarking protocols lack crucial details, which can distort evaluations.
Stefan Walk, Nikodem Majer, Konrad Schindler, Bernt Schiele
CVPR3
2010 Globally Optimal Multi-target Tracking on a Hexagonal Lattice
Anton Andriyenko, Konrad Schindler
ECCV (1)2
2010 Disparity Statistics for Pedestrian Detection: Combining Appearance, Motion and Stereo
Stefan Walk, Konrad Schindler, Bernt Schiele
ECCV (6)2
2010 Monocular 3D Scene Modeling and Inference: Understanding Multi-Object Traffic Scenes
Christian Wojek, Stefan Roth 0001, Konrad Schindler, Bernt Schiele
ECCV (4)3
2010 Multibody Structure-from-Motion in Practice
abstract
Multibody structure from motion (SfM) is the extension of classical SfM to dynamic scenes with multiple rigidly moving objects. Recent research has unveiled some of the mathematical foundations of the problem, but a practical algorithm which can handle realistic sequences is still missing. In this paper, we discuss the requirements for such an algorithm, highlight theoretical issues and practical problems, and describe how a static structure-from-motion framework needs to be extended to handle real dynamic scenes. Theoretical issues include different situations in which the number of independently moving scene objects changes: Moving objects can enter or leave the field of view, merge into the static background (e.g., when a car is parked), or split off from the background and start moving independently. Practical issues arise due to small freely moving foreground objects with few and short feature tracks. We argue that all of these difficulties need to be handled online as structure-from-motion estimation progresses, and present an exemplary solution using the framework of probabilistic model-scoring.
Kemal Egemen Ozden, Konrad Schindler, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Tracking a hand manipulating an object
abstract
We present a method for tracking a hand while it is interacting with an object. This setting is arguably the one where hand-tracking has most practical relevance, but poses significant additional challenges: strong occlusions by the object as well as self-occlusions are the norm, and classical anatomical constraints need to be softened due to the external forces between hand and object. To achieve robustness to partial occlusions, we use an individual local tracker for each segment of the articulated structure. The segments are connected in a pairwise Markov random field, which enforces the anatomical hand structure through soft constraints on the joints between adjacent segments. The most likely hand configuration is found with belief propagation. Both range and color data are used as input. Experiments are presented for synthetic data with ground truth and for real data of people manipulating objects.
Henning Hamer, Konrad Schindler, Esther Koller-Meier, Luc Van Gool
ICCV2
2009 You'll never walk alone: Modeling social behavior for multi-target tracking
abstract
Object tracking typically relies on a dynamic model to predict the object's location from its past trajectory. In crowded scenarios a strong dynamic model is particularly important, because more accurate predictions allow for smaller search regions, which greatly simplifies data association. Traditional dynamic models predict the location for each target solely based on its own history, without taking into account the remaining scene objects. Collisions are resolved only when they happen. Such an approach ignores important aspects of human behavior: people are driven by their future destination, take into account their environment, anticipate collisions, and adjust their trajectories at an early stage in order to avoid them. In this work, we introduce a model of dynamic social behavior, inspired by models developed for crowd simulation. The model is trained with videos recorded from birds-eye view at busy locations, and applied as a motion model for multi-people tracking from a vehicle-mounted camera. Experiments on real sequences show that accounting for social interactions and scene knowledge improves tracking performance, especially during occlusions.
Stefano Pellegrini, Andreas Ess, Konrad Schindler, Luc Van Gool
ICCV3
2009 Moving obstacle detection in highly dynamic scenes
abstract
We address the problem of vision-based multi-person tracking in busy pedestrian zones using a stereo rig mounted on a mobile platform. Specifically, we are interested in the application of such a system for supporting path planning algorithms in the avoidance of dynamic obstacles. The complexity of the problem calls for an integrated solution, which extracts as much visual information as possible and combines it through cognitive feedback. We propose such an approach, which jointly estimates camera position, stereo depth, object detections, and trajectories based only on visual information. The interplay between these components is represented in a graphical model. For each frame, we first estimate the ground surface together with a set of object detections. Based on these results, we then address object interactions and estimate trajectories. Finally, we employ the tracking results to predict future motion for dynamic objects and fuse this information with a static occupancy map estimated from dense stereo. The approach is experimentally evaluated on several long and challenging video sequences from busy inner-city locations recorded with different mobile setups. The results show that the proposed integration makes stable tracking and motion prediction possible, and thereby enables path planning in complex and highly dynamic scenes.
Andreas Ess, Bastian Leibe, Konrad Schindler, Luc Van Gool
ICRA3
2009 Robust Multiperson Tracking from a Mobile Platform
abstract
In this paper, we address the problem of multiperson tracking in busy pedestrian zones using a stereo rig mounted on a mobile platform. The complexity of the problem calls for an integrated solution that extracts as much visual information as possible and combines it through cognitive feedback cycles. We propose such an approach, which jointly estimates camera position, stereo depth, object detection, and tracking. The interplay between those components is represented by a graphical model. Since the model has to incorporate object-object interactions and temporal links to past frames, direct inference is intractable. We, therefore, propose a two-stage procedure: for each frame, we first solve a simplified version of the model (disregarding interactions and temporal continuity) to estimate the scene geometry and an overcomplete set of object detections. Conditioned on these results, we then address object interactions, tracking, and prediction in a second step. The approach is experimentally evaluated on several long and difficult video sequences from busy inner-city locations. Our results show that the proposed integration makes it possible to deliver robust tracking performance in scenes of realistic complexity.
Andreas Ess, Bastian Leibe, Konrad Schindler, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 A Generalisation of the ICP Algorithm for Articulated Bodies
abstract
The ICP algorithm has been extensively used in computer vision for regis-tration and tracking purposes. The original formulation of this method is restricted to the use of non-articulated models. A straightforward generali-sation to articulated structures is achievable through the joint minimisation of all the structure pose parameters, for example using Levenberg-Marquardt (LM) optimisation. However, in this approach the aligning transformation cannot be estimated in closed form, like in the original ICP, and the approach heavily suffers from local minima. To overcome this limitation, some au-thors have extended the straightforward generalisation at the cost of giving up some of the properties of ICP. In this paper, we present a generalisation of ICP to articulated structures, which preserves all the properties of the original algorithm. The key idea is to divide the articulated body into parts, which can be aligned rigidly in the way of the original ICP, with additional constraints to keep the articulated structure intact. Experiments show that our method reduces the residual registration error by a factor of ≈2. 1
Stefano Pellegrini, Konrad Schindler, Daniele Nardi
BMVC2
2008 A mobile vision system for robust multi-person tracking
abstract
We present a mobile vision system for multi-person tracking in busy environments. Specifically, the system integrates continuous visual odometry computation with tracking-by-detection in order to track pedestrians in spite of frequent occlusions and egomotion of the camera rig. To achieve reliable performance under real-world conditions, it has long been advocated to extract and combine as much visual information as possible. We propose a way to closely integrate the vision modules for visual odometry, pedestrian detection, depth estimation, and tracking. The integration naturally leads to several cognitive feedback loops between the modules. Among others, we propose a novel feedback connection from the object detector to visual odometry which utilizes the semantic knowledge of detection to stabilize localization. Feedback loops always carry the danger that erroneous feedback from one module is amplified and causes the entire system to become instable. We therefore incorporate automatic failure detection and recovery, allowing the system to continue when a module becomes unreliable. The approach is experimentally evaluated on several long and difficult video sequences from busy inner-city locations. Our results show that the proposed integration makes it possible to deliver stable tracking performance in scenes of previously infeasible complexity.
Andreas Ess, Bastian Leibe, Konrad Schindler, Luc Van Gool
CVPR3
2008 Action snippets: How many frames does human action recognition require?
abstract
Visual recognition of human actions in video clips has been an active field of research in recent years. However, most published methods either analyse an entire video and assign it a single action label, or use relatively large look-ahead to classify each frame. Contrary to these strategies, human vision proves that simple actions can be recognised almost instantaneously. In this paper, we present a system for action recognition from very short sequences (ldquosnippetsrdquo) of 1-10 frames, and systematically evaluate it on standard data sets. It turns out that even local shape and optic flow for a single frame are enough to achieve ap90% correct recognitions, and snippets of 5-7 frames (0.3-0.5 seconds of video) are enough to achieve a performance similar to the one obtainable with the entire video sequence.
Konrad Schindler, Luc Van Gool
CVPR1
2008 Articulated Multi-body Tracking under Egomotion
Stephan Gammeter, Andreas Ess, Tobias Jaeggli, Konrad Schindler, Bastian Leibe, Luc Van Gool
ECCV (2)4
2008 A Model-Selection Framework for Multibody Structure-and-Motion of Image Sequences
Konrad Schindler, David Suter, Hanzi Wang
Int. J. Comput. Vis.1
2008 Recognizing emotions expressed by body pose: A biologically inspired neural model
Konrad Schindler, Luc Van Gool, Béatrice de Gelder
Neural Networks1
2008 Coupled Object Detection and Tracking from Static Cameras and Moving Vehicles
abstract
We present a novel approach for multi-object tracking which considers object detection and spacetime trajectory estimation as a coupled optimization problem. Our approach is formulated in a Minimum Description Length hypothesis selection framework, which allows it to recover from mismatches and temporarily lost tracks. Building upon a state-of-the-art object detector, it performs multiview/multicategory object recognition to detect cars and pedestrians in the input images. The 2D object detections are checked for their consistency with (automatically estimated) scene geometry and are converted to 3D observations which are accumulated in a world coordinate frame. A subsequent trajectory estimation module analyzes the resulting 3D observations to find phyically plausible spacetime trajectories. Tracking is achieved by performing model selection after every frame. At each time instant, our approach searches for the globally optimal set of spacetime trajectories which provides the best explanation for the current image and for all evidence collected so far, while satisfying the constraints that no two objects may occupy the same physical space, nor explain the same image pixels at any point in time. Successful trajectory hypotheses are then fed back to guide object detection in future frames. The optimization procedure is kept efficient throught incremental computation and conservative hypothesis pruning. We evaluate our approach on several challenging video sequences and demonstrate its performance on both a surveillance-type scenario and a scenario where the input videos are taken from inside a moving vehicle passing through crowded city areas.
Bastian Leibe, Konrad Schindler, Nico Cornelis, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Object detection by global contour shape
Konrad Schindler, David Suter
Pattern Recognit.1
2007 Coupled Detection and Trajectory Estimation for Multi-Object Tracking
abstract
We present a novel approach for multi-object tracking which considers object detection and spacetime trajectory estimation as a coupled optimization problem. It is formulated in a hypothesis selection framework and builds upon a state-of-the-art pedestrian detector. At each time instant, it searches for the globally optimal set of spacetime trajectories which provides the best explanation for the current image and for all evidence collected so far, while satisfying the constraints that no two objects may occupy the same physical space, nor explain the same image pixels at any point in time. Successful trajectory hypotheses are fed back to guide object detection in future frames. The optimization procedure is kept efficient through incremental computation and conservative hypothesis pruning. The resulting approach can initialize automatically and track a large and varying number of persons over long periods and through complex scenes with clutter, occlusions, and large-scale background changes. Also, the global optimization framework allows our system to recover from mismatches and temporarily lost tracks. We demonstrate the feasibility of the proposed approach on several challenging video sequences.
Bastian Leibe, Konrad Schindler, Luc Van Gool
ICCV2
2007 Simultaneous Segmentation and 3D Reconstruction of Monocular Image Sequences
abstract
When trying to extract 3D scene information and camera motion from an image sequence alone, it is often necessary to cope with independently moving objects. Recent research has unveiled some of the mathematical foundations of the problem, but a general and practical algorithm, which can handle long, realistic sequences, is still missing. In this paper, we identify the necessary parts of such an algorithm, highlight both unexplored theoretical issues and practical challenges, and propose solutions. Theoretical issues include proper handling of different situations, in which the number of independent motions changes: objects can enter the scene, objects previously moving together can split and follow independent trajectories, or independently moving objects can merge into one common motion. We derive model scoring criteria to handle these changes in the number of segments. A further theoretical issue is the resolution of the relative scale ambiguity between such changes. Practical issues include robust 3D reconstruction of freely moving foreground objects, which often have few and short feature tracks. The proposed framework simultaneously tracks features, groups them into rigidly moving segments, and reconstructs all segments in 3D. Such an online approach, as opposed to batch processing techniques, which first track features, and then perform segmentation and reconstruction, is vital in order to handle small foreground objects.
Kemal Egemen Ozden, Konrad Schindler, Luc Van Gool
ICCV2
2007 Extrapolating Learned Manifolds for Human Activity Recognition
abstract
The problem of human activity recognition via visual stimuli can be approached using manifold learning, since the silhouette (binary) images of a person undergoing a smooth motion can be represented as a manifold in the image space. While manifold learning methods allow the characterization of the activity manifolds, performing activity recognition requires distinguishing between manifolds. This invariably involves the extrapolation of learned activity manifolds to new silhouettes -a task that is not fully addressed in the literature. This paper investigates and compares methods for the extrapolation of learned manifolds within the context of activity recognition. Also, the problem of obtaining dense samples for learning human silhouette manifolds is addressed.
Tat-Jun Chin, Liang Wang 0001, Konrad Schindler, David Suter
ICIP (1)3
2007 Adaptive Object Tracking Based on an Effective Appearance Filter
abstract
We propose a similarity measure based on a Spatial-color Mixture of Gaussians (SMOG) appearance model for particle filters. This improves on the popular similarity measure based on color histograms because it considers not only the colors in a region but also the spatial layout of the colors. Hence, the SMOG-based similarity measure is more discriminative. To efficiently compute the parameters for SMOG, we propose a new technique, with which the computational time is greatly reduced. We also extend our method by integrating multiple cues to increase the reliability and robustness. Experiments show that our method can successfully track objects in many difficult situations.
Hanzi Wang, David Suter, Konrad Schindler, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Smooth Foreground-Background Segmentation for Video Processing
Konrad Schindler, Hanzi Wang
ACCV (2)1
2006 Perspective n-View Multibody Structure-and-Motion Through Model Selection
Konrad Schindler, James U, Hanzi Wang
ECCV (1)1
2006 Effective Appearance Model and Similarity Measure for Particle Filtering and Visual Tracking
Hanzi Wang, David Suter, Konrad Schindler
ECCV (3)3
2006 Geometry and construction of straight lines in log-polar images
Konrad Schindler
Comput. Vis. Image Underst.1
2006 Piecewise planar scene reconstruction from sparse correspondences
Friedrich Fraundorfer, Konrad Schindler, Horst Bischof
Image Vis. Comput.2
2006 Two-View Multibody Structure-and-Motion with Outliers through Model Selection
abstract
Multibody structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: Given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task is to find an optimal set of motion models for the measurements. It is solved through Monte-Carlo sampling, careful statistical analysis of the sampled set of motion models, and simultaneous selection of multiple motion models to best explain the measurements. The framework is not restricted to any particular model selection mechanism because it is developed from a Bayesian viewpoint: Different model selection criteria are seen as different priors for the set of moving objects, which allow one to bias the selection procedure for different purposes.
Konrad Schindler, David Suter
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Two-View Multibody Structure-and-Motion with Outliers
abstract
Multi-body structure-and-motion (MSaM) is the problem to establish the multiple-view geometry of several views of a 3D scene taken at different times, where the scene consists of multiple rigid objects moving relative to each other. We examine the case of two views. The setting is the following: given are a set of corresponding image points in two images, which originate from an unknown number of moving scene objects, each giving rise to a motion model. Furthermore, the measurement noise is unknown, and there are a number of gross errors, which are outliers to all models. The task to find an optimal set of motion models for the measurements is solved through Monte-Carlo sampling, careful statistical analysis of the data and simultaneous selection of multiple motion models.
Konrad Schindler, David Suter
CVPR (2)1
2005 Spatially consistent 3D motion segmentation
abstract
3D motion segmentation is the task to cluster corresponding points in multiple (at least two) images, so that each cluster corresponds to a 3D motion in the underlying 3D scene. The problem can be divided into two stages: first, all motion models required to describe the scene have to be found. Second, each correspondence has to be assigned to the correct model. This paper is concerned with the second part. A natural procedure is to assign each correspondence to a motion, such that the a-posteriori likelihood of the description is maximized. However, this is not trivial, since the likelihoods of different correspondences are not independent: neighboring correspondences tend to belong to the same motion, a fact commonly referred to as "smoothness" or "spatial consistency". To account for this fact, we model the set of multiview correspondences as an irregular Markov random field (MRF). The MRF is then optimized with recent graph-based methods, and individual clique potentials are inspected for fine-grained outlier detection.
Konrad Schindler
ICIP (3)1
2002 Metropogis: a city information system
abstract
We report on a new system to augment a 3D block model of a real city obtained from aerial photogrammetry or aerial laser scanning with geo-referenced terrestrial data of the facades. The terrestrial images are acquired by a hand-held digital consumer camera. The relative orientation of the photographs is calculated automatically and fitted towards the 3D block model with minimized human input using vanishing points. The extraction of 3D primitives on the facades is based on line matching over multiple oriented images. The introduced city information system delivers a fully 3D geographic information data set and is called MetropoGIS.
Andreas Klaus, Joachim Bauer, Konrad F. Karner, Konrad Schindler
ICIP (3)4