VLDB 2026 Research / reviewers in the wild / expert
Jan Dirk Wegner
dblp:66/8991 · also Jan D. Wegner
· DBLP profile ↗
31ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0002-0290-6901ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Point2CAD: Reverse Engineering CAD Models from 3D Point CloudsabstractComputer-Aided Design (CAD) model reconstruction from point clouds is an important problem at the intersection of computer vision, graphics, and machine learning; it saves the designer significant time when iterating on in-the-wild objects. Recent advancements in this direction achieve relatively reliable semantic segmentation but still struggle to produce an adequate topology of the CAD model. In this work, we analyze the current state of the art for that ill-posed task and identify shortcomings of existing methods. We propose a hybrid analyticneural reconstruction scheme that bridges the gap between segmented point clouds and structured CAD models and can be readily combined with different segmentation backbones. Moreover, to power the surface fitting stage, we propose a novel implicit neural representation of freeform surfaces, driving up the performance of our overall CAD reconstruction scheme. We extensively evaluate our method on the popular ABC benchmark of CAD models and set a new state-of-the-art for that dataset. Code is available at https://github.com/YujiaLiu76/point2cad. Yujia Liu 0002, Anton Obukhov, Jan Dirk Wegner, Konrad Schindler |
CVPR | 3 |
| 2024 | TetraDiffusion: Tetrahedral Diffusion Models for 3D Shape Generation
Nikolai Kalischek, Torben Peters, Jan Dirk Wegner, Konrad Schindler |
ECCV (53) | 3 |
| 2024 | Recognition of Unseen Bird Species by Learning from Field GuidesabstractWe exploit field guides to learn bird species recognition, in particular zero-shot recognition of unseen species. Illustrations contained in field guides deliberately focus on discriminative properties of each species, and can serve as side information to transfer knowledge from seen to unseen bird species. We study two approaches: (1) a contrastive encoding of illustrations, which can be fed into standard zero-shot learning schemes; and (2) a novel method that leverages the fact that illustrations are also images and as such structurally more similar to photographs than other kinds of side information. Our results show that illustrations from field guides, which are readily available for a wide range of species, are indeed a competitive source of side information for zero-shot learning. On a subset of the iNaturalist2021 dataset with 749 seen and 739 unseen species, we obtain a classification accuracy of unseen bird species of 12% @top-1 and 38% @top-10, which shows the potential of field guides for challenging real-world scenarios with many species. Our code is available at https://github.com/ac-rodriguez/zsl_billow. Andrés C. Rodríguez, Stefano D'Aronco, Rodrigo Caye Daudt, Jan Dirk Wegner, Konrad Schindler |
WACV | 4 |
| 2023 | BiasBed - Rigorous Texture Bias EvaluationabstractThe well-documented presence of texture bias in modern convolutional neural networks has led to a plethora of algorithms that promote an emphasis on shape cues, often to support generalization to new domains. Yet, common datasets, benchmarks and general model selection strategies are missing, and there is no agreed, rigorous evaluation protocol. In this paper, we investigate difficulties and limitations when training networks with reduced texture bias. In particular, we also show that proper evaluation and meaningful comparisons between methods are not trivial. We introduce BiasBed, a testbed for texture- and style-biased training, including multiple datasets and a range of existing algorithms. It comes with an extensive evaluation protocol that includes rigorous hypothesis testing to gauge the significance of the results, despite the considerable training instability of some style bias methods. Our extensive experiments, shed new light on the need for careful, statistically founded evaluation protocols for style bias (and beyond). E.g., we find that some algorithms proposed in the literature do not significantly mitigate the impact of style bias at all. With the release of BiasBed, we hope to foster a common understanding of consistent and meaningful comparisons, and consequently faster progress towards learning methods free of texture bias. Code is available at https://github.com/D1noFuzi/BiasBed Nikolai Kalischek, Rodrigo Caye Daudt, Torben Peters, Reinhard Furrer, Jan Dirk Wegner, Konrad Schindler |
CVPR | 5 |
| 2023 | Fine-Grained Species Recognition With Privileged Pooling: Better Sample Efficiency Through Supervised AttentionabstractWe propose a scheme for supervised image classification that uses privileged information, in the form of keypoint annotations for the training data, to learn strong models from small and/or biased training sets. Our main motivation is the recognition of animal species for ecological applications such as biodiversity modelling, which is challenging because of long-tailed species distributions due to rare species, and strong dataset biases such as repetitive scene background in camera traps. To counteract these challenges, we propose a visual attention mechanism that is supervised via keypoint annotations that highlight important object parts. This privileged information, implemented as a novel privileged pooling operation, is only required during training and helps the model to focus on regions that are discriminative. In experiments with three different animal species datasets, we show that deep networks with privileged pooling can use small training sets more efficiently and generalize better. Andrés C. Rodríguez, Stefano D'Aronco, Konrad Schindler, Jan Dirk Wegner |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Learning Graph Regularisation for Guided Super-ResolutionabstractWe introduce a novel formulation for guided super-resolution. Its core is a differentiable optimisation layer that operates on a learned affinity graph. The learned graph potentials make it possible to leverage rich contextual information from the guide image, while the explicit graph optimisation within the architecture guarantees rigorous fidelity of the high-resolution target to the low-resolution source. With the decision to employ the source as a constraint rather than only as an input to the prediction, our method differs from state-of-the-art deep architectures for guided super-resolution, which produce targets that, when downsampled, will only approximately reproduce the source. This is not only theoretically appealing, but also produces crisper, more natural-looking images. A key property of our method is that, although the graph connectivity is restricted to the pixel lattice, the associated edge potentials are learned with a deep feature extractor and can encode rich context information over large receptive fields. By taking advantage of the sparse graph connectivity, it becomes possible to propagate gradients through the optimisation layer and learn the edge potentials from data. We extensively evaluate our method on several datasets, and consistently outperform recent baselines in terms of quantitative reconstruction errors, while also delivering visually sharper outputs. Moreover, we demonstrate that our method generalises particularly well to new datasets not seen during training. Riccardo de Lutio, Alexander Becker 0002, Stefano D'Aronco, Stefania Russo, Jan Dirk Wegner, Konrad Schindler |
CVPR | 5 |
| 2022 | FiLM-Ensemble: Probabilistic Deep Learning via Feature-wise Linear ModulationabstractThe ability to estimate epistemic uncertainty is often crucial when deploying machine learning in the real world, but modern methods often produce overconfident, uncalibrated uncertainty predictions. A common approach to quantify epistemic uncertainty, usable across a wide class of prediction models, is to train a model ensemble. In a naive implementation, the ensemble approach has high computational cost and high memory demand. This challenges in particular modern deep learning, where even a single deep network is already demanding in terms of compute and memory, and has given rise to a number of attempts to emulate the model ensemble without actually instantiating separate ensemble members. We introduce FiLM-Ensemble, a deep, implicit ensemble method based on the concept of Feature-wise Linear Modulation (FiLM). That technique was originally developed for multi-task learning, with the aim of decoupling different tasks. We show that the idea can be extended to uncertainty quantification: by modulating the network activations of a single deep network with FiLM, one obtains a model ensemble with high diversity, and consequently well-calibrated estimates of epistemic uncertainty, with low computational overhead in comparison. Empirically, FiLM-Ensemble outperforms other implicit ensemble methods, and it comes very close to the upper bound of an explicit ensemble of networks (sometimes even beating it), at a fraction of the memory cost. Mehmet Ozgur Turkoglu, Alexander Becker 0002, Hüseyin Anil Gündüz, Mina Rezaei, Bernd Bischl, Rodrigo Caye Daudt, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler |
NeurIPS | 8 |
| 2022 | Gating Revisited: Deep Multi-Layer RNNs That can be TrainedabstractWe propose a new STAckable Recurrent cell (STAR) for recurrent neural networks (RNNs), which has fewer parameters than widely used LSTM [16] and GRU [10] while being more robust against vanishing or exploding gradients. Stacking recurrent units into deep architectures suffers from two major limitations: (i) many recurrent cells (e.g., LSTMs) are costly in terms of parameters and computation resources; and (ii) deep RNNs are prone to vanishing or exploding gradients during training. We investigate the training of multi-layer RNNs and examine the magnitude of the gradients as they propagate through the network in the "vertical" direction. We show that, depending on the structure of the basic recurrent unit, the gradients are systematically attenuated or amplified. Based on our analysis we design a new type of gated cell that better preserves gradient magnitude. We validate our design on a large number of sequence modelling tasks and demonstrate that the proposed STAR cell allows to build and train deeper recurrent architectures, ultimately leading to improved performance while being computationally more efficient. Mehmet Ozgur Turkoglu, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Crop Classification Under Varying Cloud Cover With Neural Ordinary Differential EquationsabstractOptical satellite sensors cannot see the earth’s surface through clouds. Despite the periodic revisit cycle, image sequences acquired by earth observation satellites are, therefore,irregularlysampled in time. State-of-the-art methods for crop classification (and other time-series analysis tasks) rely on techniques that implicitly assume regular temporal spacing between observations, such as recurrent neural networks (RNNs). We propose to use neural ordinary differential equations (NODEs) in combination with RNNs to classify crop types in irregularly spaced image sequences. The resulting ODE-RNN models consist of two steps: an update step, where a recurrent unit assimilates new input data into the model’s hidden state, and a prediction step, in which NODE propagates the hidden state until the next observation arrives. The prediction step is based on a continuous representation of the latent dynamics, which has several advantages. At the conceptual level, it is a more natural way to describe the mechanisms that govern the phenological cycle. From a practical point of view, it makes it possible to sample the system state at arbitrary points in time such that one can integrate observations whenever they are available and extrapolate beyond the last observation. Our experiments show that ODE-RNN, indeed, improves classification accuracy over common baselines, such as LSTM, GRU, temporal convolutional network, and transformer. The gains are most prominent in the challenging scenario where only few observations are available (i.e., frequent cloud cover). Moreover, we show that the ability to extrapolate translates to better classification performance early in the season, which is important for forecasting. Nando Metzger, Mehmet Ozgur Turkoglu, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | In the Light of Feature Distributions: Moment Matching for Neural Style TransferabstractStyle transfer aims to render the content of a given image in the graphical/artistic style of another image. The fundamental concept underlying Neural Style Transfer (NST) is to interpret style as a distribution in the feature space of a Convolutional Neural Network, such that a desired style can be achieved by matching its feature distribution. We show that most current implementations of that concept have important theoretical and practical limitations, as they only partially align the feature distributions. We propose a novel approach that matches the distributions more precisely, thus reproducing the desired style more faithfully, while still being computationally efficient. Specifically, we adapt the dual form of Central Moment Discrepancy (CMD), as recently proposed for domain adaptation, to minimize the difference between the target style and the feature distribution of the output image. The dual interpretation of this metric explicitly matches all higher-order centralized moments and is therefore a natural extension of existing NST methods that only take into account the first and second moments. Our experiments confirm that the strong theoretical properties also translate to visually better style transfer, and better disentangle style from semantic image content. Nikolai Kalischek, Jan Dirk Wegner, Konrad Schindler |
CVPR | 2 |
| 2021 | PC2WF: 3D Wireframe Reconstruction from Raw Point Clouds
Yujia Liu 0002, Stefano D'Aronco, Konrad Schindler, Jan Dirk Wegner |
ICLR | 4 |
| 2020 | Learning Multiview 3D Point Cloud RegistrationabstractWe present a novel, end-to-end learnable, multiview 3D point cloud registration algorithm. Registration of multiple scans typically follows a two-stage pipeline: the initial pairwise alignment and the globally consistent refinement. The former is often ambiguous due to the low overlap of neighboring point clouds, symmetries and repetitive scene parts. Therefore, the latter global refinement aims at establishing the cyclic consistency across multiple scans and helps in resolving the ambiguous cases. In this paper we propose, to the best of our knowledge, the first end-to-end algorithm for joint learning of both parts of this two-stage problem. Experimental evaluation on well accepted benchmark datasets shows that our approach outperforms the state-of-the-art by a significant margin, while being end-to-end trainable and computationally less costly. Moreover, we present detailed analysis and an ablation study that validate the novel components of our approach. The source code and pretrained models are publicly available under https://github.com/zgojcic/3D_multiview_reg. Zan Gojcic, Caifa Zhou, Jan Dirk Wegner, Leonidas J. Guibas, Tolga Birdal |
CVPR | 3 |
| 2020 | GeoGraph: Graph-Based Multi-view Object Detection with Geometric Cues End-to-End
Ahmed Samy Nassar, Stefano D'Aronco, Sébastien Lefèvre, Jan Dirk Wegner |
ECCV (7) | 4 |
| 2020 | HistoNet: Predicting size histograms of object instancesabstractWe propose to predict histograms of object sizes in crowded scenes directly without any explicit object instance segmentation. What makes this task challenging is the high density of objects (of the same category), which makes instance identification hard. Instead of explicitly segmenting object instances, we show that directly learning histograms of object sizes improves accuracy while using drastically less parameters. This is very useful for application scenarios where explicit, pixel-accurate instance segmentation is not needed, but there lies interest in the overall distribution of instance sizes. Our core applications are in biology, where we estimate the size distribution of soldier fly larvae, and medicine, where we estimate the size distribution of cancer cells as an intermediate step to calculate the tumor cellularity score. Given an image with hundreds of small object instances, we output the total count and the size histogram. We also provide a new data set for this task, the FlyLarvae data set, which consists of 11,000 larvae instances labeled pixel-wise. Our method results in an overall improvement in the count and size distribution prediction as compared to state-of-the-art instance segmentation method Mask R-CNN [11]. Kishan Sharma, Moritz Gold, Christian Zurbruegg, Laura Leal-Taixé, Jan Dirk Wegner |
WACV | 5 |
| 2020 | Inference, Learning and Attention Mechanisms that Exploit and Preserve Sparsity in CNNs
Timo Hackel, Mikhail Usvyatsov, Silvano Galliani, Jan Dirk Wegner, Konrad Schindler |
Int. J. Comput. Vis. | 4 |
| 2019 | The Perfect Match: 3D Point Cloud Matching With Smoothed DensitiesabstractWe propose 3DSmoothNet, a full workflow to match 3D point clouds with a siamese deep learning architecture and fully convolutional layers using a voxelized smoothed density value (SDV) representation. The latter is computed per interest point and aligned to the local reference frame (LRF) to achieve rotation invariance. Our compact, learned, rotation invariant 3D point cloud descriptor achieves 94.9% average recall on the 3DMatch benchmark data set, outperforming the state-of-the-art by more than 20 percent points with only 32 output dimensions. This very low output dimension allows for near realtime correspondence search with 0.1 ms per feature point on a standard PC. Our approach is sensor- and scene-agnostic because of SDV, LRF and learning highly descriptive features with fully convolutional layers. We show that 3DSmoothNet trained only on RGB-D indoor scenes of buildings achieves 79.0% average recall on laser scans of outdoor vegetation, more than double the performance of our closest, learning-based competitors. Code, data and pre-trained models are available online at https://github.com/zgojcic/3DSmoothNet. Zan Gojcic, Caifa Zhou, Jan Dirk Wegner, Andreas Wieser |
CVPR | 3 |
| 2019 | Topological Map Extraction From Overhead ImagesabstractWe propose a new approach, named PolyMapper, to circumvent the conventional pixel-wise segmentation of (aerial) images and predict objects in a vector representation directly. PolyMapper directly extracts the topological map of a city from overhead images as collections of building footprints and road networks. In order to unify the shape representation for different types of objects, we also propose a novel sequentialization method that reformulates a graph structure as closed polygons. Experiments are conducted on both existing and self-collected large-scale datasets of several cities. Our empirical results demonstrate that our end-to-end learnable model is capable of drawing polygons of building footprints and road networks that very closely approximate the structure of existing online map services, in a fully automated manner. Quantitative and qualitative comparison to the state-of-the-arts also show that our approach achieves good levels of performance. To the best of our knowledge, the automatic extraction of large-scale topological maps is a novel contribution in the remote sensing community that we believe will help develop models with more informed geometrical constraints. Zuoyue Li, Jan Dirk Wegner, Aurélien Lucchi |
ICCV | 2 |
| 2019 | Guided Super-Resolution As Pixel-to-Pixel TransformationabstractGuided super-resolution is a unifying framework for several computer vision tasks where the inputs are a low-resolution source image of some target quantity (e.g., perspective depth acquired with a time-of-flight camera) and a high-resolution guide image from a different domain (e.g., a grey-scale image from a conventional camera); and the target output is a high-resolution version of the source (in our example, a high-res depth map). The standard way of looking at this problem is to formulate it as a super-resolution task, i.e., the source image is upsampled to the target resolution, while transferring the missing high-frequency details from the guide. Here, we propose to turn that interpretation on its head and instead see it as a pixel-to-pixel mapping of the guide image to the domain of the source image. The pixel-wise mapping is parametrised as a multi-layer perceptron, whose weights are learned by minimising the discrepancies between the source image and the downsampled target image. Importantly, our formulation makes it possible to regularise only the mapping function, while avoiding regularisation of the outputs; thus producing crisp, natural-looking images. The proposed method is unsupervised, using only the specific source and guide images to fit the mapping. We evaluate our method on two different tasks, super-resolution of depth maps and of tree height maps. In both cases, we clearly outperform recent baselines in quantitative comparisons, while delivering visually much sharper outputs. Riccardo de Lutio, Stefano D'Aronco, Jan Dirk Wegner, Konrad Schindler |
ICCV | 3 |
| 2019 | Simultaneous Multi-View Instance Detection With Learned Geometric Soft-ConstraintsabstractWe propose to jointly learn multi-view geometry and warping between views of the same object instances for robust cross-view object detection. What makes multi-view object instance detection difficult are strong changes in viewpoint, lighting conditions, high similarity of neighbouring objects, and strong variability in scale. By turning object detection and instance re-identification in different views into a joint learning task, we are able to incorporate both image appearance and geometric soft constraints into a single, multi-view detection process that is learnable end-to-end. We validate our method on a new, large data set of street-level panoramas of urban objects and show superior performance compared to various baselines. Our contribution is threefold: a large-scale, publicly available data set for multi-view instance detection and re-identification; an annotation tool custom-tailored for multi-view instance detection; and a novel, holistic multi-view instance detection and re-identification method that jointly models geometry and appearance across views. Ahmed Samy Nassar, Sébastien Lefèvre, Jan Dirk Wegner |
ICCV | 3 |
| 2017 | Semantically Informed Multiview Surface RefinementabstractWe present a method to jointly refine the geometry and semantic segmentation of 3D surface meshes. Our method alternates between updating the shape and the semantic labels. In the geometry refinement step, the mesh is deformed with variational energy minimization, such that it simultaneously maximizes photo-consistency and the compatibility of the semantic segmentations across a set of calibrated images. Label-specific shape priors account for interactions between the geometry and the semantic labels in 3D. In the semantic segmentation step, the labels on the mesh are updated with MRF inference, such that they are compatible with the semantic segmentations in the input images. Also, this step includes prior assumptions about the surface shape of different semantic classes. The priors induce a tight coupling, where semantic information influences the shape update and vice versa. Specifically, we introduce priors that favor (i) adaptive smoothing, depending on the class label; (ii) straightness of class boundaries; and (iii) semantic labels that are consistent with the surface orientation. The novel mesh-based reconstruction is evaluated in a series of experiments with real and synthetic data. We compare both to state-of-the-art, voxel-based semantic 3D reconstruction, and to purely geometric mesh refinement, and demonstrate that the proposed scheme yields improved 3D geometry as well as an improved semantic segmentation. Maros Blaha, Mathias Rothermel, Martin R. Oswald, Torsten Sattler, Audrey Richard, Jan Dirk Wegner, Marc Pollefeys, Konrad Schindler |
ICCV | 6 |
| 2017 | Semantic segmentation of aerial images with explicit class-boundary modelingabstractIn this work we propose an end-to-end trainable supervised Deep Convolutional Neural Network (DCNN) targeting the task of semantic-segmentation with the addition of class-aware boundary detection. Through this explicit modeling of the class-boundaries, we enforce the network to extract coherent and complete objects, suppressing the uncertainty influencing these regions. Importantly, we show that class-boundary networks in conjunction with DCNN performs optimally, achieving over 90% overall accuracy (OA) on the challenging ISPRS Vaihingen Semantic Segmentation benchmark. Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner, Mihai Datcu, Uwe Stilla |
IGARSS | 3 |
| 2017 | Toward Seamless Multiview Scene Analysis From Satellite to Street LevelabstractIn this paper, we discuss and review how combined multiview imagery from satellite to street level can benefit scene analysis. Numerous works exist that merge information from remote sensing and images acquired from the ground for tasks such as object detection, robots guidance, or scene understanding. What makes the combination of overhead and street-level images challenging are the strongly varying viewpoints, the different scales of the images, their illuminations and sensor modality, and time of acquisition. Direct (dense) matching of images on a per-pixel basis is thus often impossible, and one has to resort to alternative strategies that will be discussed in this paper. For such purpose, we review recent works that attempt to combine images taken from the ground and overhead views for purposes like scene registration, reconstruction, or classification. After the theoretical review, we present three recent methods to showcase the interest and potential impact of such fusion on real applications (change detection, image orientation, and tree cataloging), whose logic can then be reused to extend the use of ground-based images in remote sensing andvice versa. Through this review, we advocate that cross fertilization between remote sensing, computer vision, and machine learning is very valuable to make the best of geographic data available from Earth observation sensors and ground imagery. Despite its challenges, we believe that integrating these complementary data sources will lead to major breakthroughs in Big GeoData. It will open new perspectives for this exciting and emerging field. Sébastien Lefèvre, Devis Tuia, Jan Dirk Wegner, Timothée Produit, Ahmed Samy Nassar |
Proc. IEEE | 3 |
| 2017 | Learning Aerial Image Segmentation From Online MapsabstractThis paper deals with semantic segmentation of high-resolution (aerial) images where a semantic class label is assigned to each pixel via supervised classification as a basis for automatic map generation. Recently, deep convolutional neural networks (CNNs) have shown impressive performance and have quickly become the de-facto standard for semantic segmentation, with the added benefit that task-specific feature design is no longer necessary. However, a major downside of deep learning methods is that they are extremely data hungry, thus aggravating the perennial bottleneck of supervised classification, to obtain enough annotated training data. On the other hand, it has been observed that they are rather robust against noise in the training labels. This opens up the intriguing possibility to avoid annotating huge amounts of training data, and instead train the classifier from existing legacy data or crowd-sourced maps that can exhibit high levels of noise. The question addressed in this paper is: can training with large-scale publicly available labels replace a substantial part of the manual labeling effort and still achieve sufficient performance? Such data will inevitably contain a significant portion of errors, but in return virtually unlimited quantities of it are available in larger parts of the world. We adapt a state-of-the-art CNN architecture for semantic segmentation of buildings and roads in aerial images, and compare its performance when using different training data sets, ranging from manually labeled pixel-accurate ground truth of the same city to automatic training data derived from OpenStreetMap data from distant locations. We report our results that indicate that satisfying performance can be obtained with significantly less manual annotation effort, by exploiting noisy large-scale training data. Pascal Kaiser, Jan Dirk Wegner, Aurélien Lucchi, Martin Jaggi, Thomas Hofmann 0001, Konrad Schindler |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Large-Scale Semantic 3D Reconstruction: An Adaptive Multi-resolution Model for Multi-class Volumetric LabelingabstractWe propose an adaptive multi-resolution formulation of semantic 3D reconstruction. Given a set of images of a scene, semantic 3D reconstruction aims to densely reconstruct both the 3D shape of the scene and a segmentation into semantic object classes. Jointly reasoning about shape and class allows one to take into account class-specific shape priors (e.g., building walls should be smooth and vertical, and vice versa smooth, vertical surfaces are likely to be building walls), leading to improved reconstruction results. So far, semantic 3D reconstruction methods have been limited to small scenes and low resolution, because of their large memory footprint and computational cost. To scale them up to large scenes, we propose a hierarchical scheme which refines the reconstruction only in regions that are likely to contain a surface, exploiting the fact that both high spatial resolution and high numerical precision are only required in those regions. Our scheme amounts to solving a sequence of convex optimizations while progressively removing constraints, in such a way that the energy, in each iteration, is the tightest possible approximation of the underlying energy at full resolution. In our experiments the method saves up to 98% memory and 95% computation time, without any loss of accuracy. Maros Blaha, Christoph Vogel, Audrey Richard, Jan Dirk Wegner, Thomas Pock, Konrad Schindler |
CVPR | 4 |
| 2016 | Contour Detection in Unstructured 3D Point CloudsabstractWe describe a method to automatically detect contours, i.e. lines along which the surface orientation sharply changes, in large-scale outdoor point clouds. Contours are important intermediate features for structuring point clouds and converting them into high-quality surface or solid models, and are extensively used in graphics and mapping applications. Yet, detecting them in unstructured, inhomogeneous point clouds turns out to be surprisingly difficult, and existing line detection algorithms largely fail. We approach contour extraction as a two-stage discriminative learning problem. In the first stage, a contour score for each individual point is predicted with a binary classifier, using a set of features extracted from the point's neighborhood. The contour scores serve as a basis to construct an overcomplete graph of candidate contours. The second stage selects an optimal set of contours from the candidates. This amounts to a further binary classification in a higher-order MRF, whose cliques encode a preference for connected contours and penalize loose ends. The method can handle point clouds > 107 points in a couple of minutes, and vastly outperforms a baseline that performs Canny-style edge detection on a range image representation of the point cloud. Timo Hackel, Jan Dirk Wegner, Konrad Schindler |
CVPR | 2 |
| 2016 | Cataloging Public Objects Using Aerial and Street-Level Images - Urban TreesabstractEach corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design computer vision systems that will help us search, catalog and monitor public infrastructure, buildings and artifacts. We explore the architecture and feasibility of such a system. The main technical challenge is combining test time information from multiple views of each geographic location (e.g., aerial and street views). We implement two modules: det2geo, which detects the set of locations of objects belonging to a given category, and geo2cat, which computes the fine-grained category of the object at a given location. We introduce a solution that adapts state-of the-art CNN-based object detectors and classifiers. We test our method on "Pasadena Urban Trees", a new dataset of 80,000 trees with geographic and species annotations, and show that combining multiple views significantly improves both tree detection and tree species classification, rivaling human performance. Jan Dirk Wegner, Steve Branson, David Hall 0002, Konrad Schindler, Pietro Perona |
CVPR | 1 |
| 2015 | Features, Color Spaces, and Boosting: New Insights on Semantic Classification of Remote Sensing ImagesabstractA major yet largely unsolved problem in the semantic classification of very high resolution remote sensing images is the design and selection of appropriate features. At a ground sampling distance below half a meter, fine-grained texture details of objects emerge and lead to a large intraclass variability while generally keeping the between-class variability at a low level. Usually, the user makes an educated guess on what features seem to appropriately capture characteristic object class patterns. Here, we propose to avoid manual feature selection and let a boosting classifier choose optimal features from a vast Randomized Quasi-Exhaustive (RQE) set of feature candidates directly during training. This RQE feature set consists of a multitude of very simple features that are computed efficiently via integral images inside a sliding window. This simple but comprehensive feature candidate set enables the boosting classifier to assemble the most discriminative textures at different scale levels to classify a small number of broad urban land-cover classes. We do an extensive evaluation on several data sets and compare performance against multiple feature extraction baselines in different color spaces. In addition, we verify experimentally if we gain any classification accuracy if moving from boosting stumps to trees. Cross-validation minimizes the possible bias caused by specific training/testing setups. It turns out that boosting in combination with the proposed RQE feature set outperforms all baseline features while still remaining computationally efficient. Particularly boosting trees (instead of stumps) captures class patterns so well that results suggest to completely leave feature selection to the classifier. Piotr Tokarczyk, Jan Dirk Wegner, Stefan Walk, Konrad Schindler |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Combining High-Resolution Optical and InSAR Features for Height Estimation of Buildings With Flat RoofsabstractIn this paper, we contribute to answer the question: How accurately can we estimate heights of buildings with flat roofs given one high-resolution single-pass interferometric synthetic aperture radar (InSAR) image pair and one aerial orthophoto? What makes this problem challenging are the different sensor geometries and the sound stochastic combination of all available elevation cues. We revisit already existing methods and develop novel approaches to determine building heights. A rigorous stochastic approach based on least squares adjustment with functionally dependent parameters is introduced to combine all height measurements per building to one robust height estimate. Observation accuracies of the stochastic model are either taken from the literature or estimated empirically. A major benefit of adjustment is that it delivers a posterior standard deviation per height, which can be interpreted as a precision indicator and is of high relevance for practical applications. Estimated heights of an urban scene are compared to ground truth acquired with airborne laser scanning, allowing us to assess height accuracies that can be achieved under nearly optimal conditions. We conduct statistical tests that validate our model and show that a weighted combination of optical and synthetic aperture radar (SAR) data with least squares adjustment delivers reliable height estimates with meter accuracy for flat-roofed buildings. Additionally, we empirically estimate a confidence interval of the estimated heights that directly tells the user the security margin to be included, for example, in case of building evacuations for an anticipated flooding event, under the condition that the data and model have the same specifications as in this paper. Jan Dirk Wegner, Jens R. Ziehn, Uwe Sörgel |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | A Higher-Order CRF Model for Road Network ExtractionabstractThe aim of this work is to extract the road network from aerial images. What makes the problem challenging is the complex structure of the prior: roads form a connected network of smooth, thin segments which meet at junctions and crossings. This type of a-priori knowledge is more difficult to turn into a tractable model than standard smoothness or co-occurrence assumptions. We develop a novel CRF formulation for road labeling, in which the prior is represented by higher-order cliques that connect sets of super pixels along straight line segments. These long-range cliques have asymmetric PN-potentials, which express a preference to assign all rather than just some of their constituent super pixels to the road class. Thus, the road likelihood is amplified for thin chains of super pixels, while the CRF is still amenable to optimization with graph cuts. Since the number of such cliques of arbitrary length is huge, we furthermore propose a sampling scheme which concentrates on those cliques which are most relevant for the optimization. In experiments on two different databases the model significantly improves both the per-pixel accuracy and the topological correctness of the extracted roads, and outperforms both a simple smoothness prior and heuristic rule-based road completion. Jan Dirk Wegner, Javier A. Montoya-Zegarra, Konrad Schindler |
CVPR | 1 |
| 2011 | A contextual approach to building extraction combining highre-solution InSAR and optical dataabstractUrban areas are highly complex scenes and thus single buildings are often hard to extract. A combination of complementary features of high-resolution interferometric SAR (InSAR) data and optical imagery can add valuable information in case one data source leads to ambiguous results. In addition, contextual information like the sun shadow, front yards, and driveways can act as hints for buildings. We propose a Conditional Random Field (CRF) approach set up on super-pixels of an initial multi-scale segmentation to learn and infer those contextual features in addition to direct building features. We compare our results to CRFs with square image patches and CRFs with single scale super-pixels. Experiments on airborne InSAR data and optical data show that the proposed CRF with multi-scale super-pixels performs best. Jan Dirk Wegner, Uwe Sörgel |
IGARSS | 1 |
| 2010 | Building detection and height estimation from high-resolution insar and optical dataabstractState-of-the-art satellite SAR sensors acquire data of one meter geometric ground resolution, airborne sensors achieve even higher resolutions. Nonetheless, layover and occlusion hamper interpretability of such data particularly in urban scenes. In order to overcome this drawback, we use additional information from aerial photos to detect buildings. Features are extracted from both data sets and introduced to a common feature vector followed by a classification into building sites and non-building sites with Conditional Random Fields (CRF). Furthermore, we show that the different sensor geometries of the SAR and the optical sensor may be used to estimate building heights. Jan Dirk Wegner, Jens R. Ziehn, Uwe Sörgel |
IGARSS | 1 |