Uwe Stilla

dblp:s/UweStilla · DBLP profile ↗
← Back
51ranked-venue papers
2as first author
8since 2021 · last 2023
0000-0002-1184-0924ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 CASSPR: Cross Attention Single Scan Place Recognition
abstract
Place recognition based on point clouds (LiDAR) is an important component for autonomous robots or self-driving vehicles. Current SOTA performance is achieved on accumulated LiDAR submaps using either point-based or voxel-based structures. While voxel-based approaches nicely integrate spatial context across multiple scales, they do not exhibit the local precision of point-based methods. As a result, existing methods struggle with fine-grained matching of subtle geometric features in sparse single-shot Li-DAR scans. To overcome these limitations, we propose CASSPR as a method to fuse point-based and voxel-based approaches using cross attention transformers. CASSPR leverages a sparse voxel branch for extracting and aggregating information at lower resolution and a point-wise branch for obtaining fine-grained local information. CASSPR uses queries from one branch to try to match structures in the other branch, ensuring that both extract self-contained descriptors of the point cloud (rather than one branch dominating), but using both to inform the out-put global descriptor of the point cloud. Extensive experiments show that CASSPR surpasses the state-of-the-art by a large margin on several datasets (Oxford RobotCar, TUM, USyd). For instance, it achieves AR@1 of 85.6% on the TUM dataset, surpassing the strongest prior model by ~15%. Our code is publicly available.1
Yan Xia 0003, Mariia Gladkova, Rui Wang 0037, Qianyun Li, Uwe Stilla, João F. Henriques, Daniel Cremers
ICCV5
2023 SegTrans: Semantic Segmentation With Transfer Learning for MLS Point Clouds
abstract
3D point cloud semantic segmentation plays an essential role in fine-grained scene understanding from photogrammetry to autonomous driving. Although recent efforts have been made to push the 3D semantic segmentation forward, many solutions cannot generalize well to new data with different sensor configurations. For example, when transferring the segmentation model learned from terrestrial laser scanning (TLS) data to mobile laser scanning (MLS) data, the performance drops dramatically. Besides, rich-labeled data is usually required. However, labeling point cloud data is time-consuming and label-intensive in practice. In light of this, we propose SegTrans, an unsupervised domain adaption method for the point cloud semantic segmentation task, which largely improves the generalization performance from one labeled dataset (source domain) to another unlabeled dataset (target domain). Specifically, we first introduce a data selection module (DSM) to tackle the discrepancy between different datasets at the data level. Then an adversarial learning module (ALM) with an adversarial loss is iteratively implemented to align the domain-specific feature in both the source and target domains, which only consists of two fully connected layers. Experiments show the overall accuracy of the proposed method achieves 88% OA on the TUM City Campus dataset (MLS dataset) when trained on the Semantic3D dataset (TLS dataset).
Shuo Shen 0003, Yan Xia 0003, Andreas Eich, Yusheng Xu, Bisheng Yang, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.6
2023 A Lightweight and Detector-Free 3D Single Object Tracker on Point Clouds
abstract
Recent works on 3D single object tracking treat the task as a target-specific 3D detection task, where an off-the-shelf 3D detector is commonly employed for the tracking. However, it is non-trivial to perform accurate target-specific detection since the point cloud of objects in raw LiDAR scans is usually sparse and incomplete. In this paper, we address this issue by explicitly leveraging temporal motion cues and propose DMT, a Detector-free Motion-prediction-based 3D Tracking network that completely removes the usage of complicated 3D detectors and is lighter, faster, and more accurate than previous trackers. Specifically, the motion prediction module is first introduced to estimate a potential target center of the current frame in a point-cloud-free manner. Then, an explicit voting module is proposed to directly regress the 3D box from the estimated target center. Extensive experiments on KITTI and NuScenes datasets demonstrate that our DMT can still achieve better performance ($\sim $10% improvement over the NuScenes dataset) and a faster tracking speed (i.e., 72 FPS) than state-of-the-art approaches without applying any complicated 3D detectors. Our code is released athttps://github.com/jimmy-dq/DMT.
Yan Xia 0003, Qiangqiang Wu, Wei Li 0111, Antoni B. Chan, Uwe Stilla
IEEE Trans. Intell. Transp. Syst.5
2022 Pairwise Point Cloud Registration Using Graph Matching and Rotation-Invariant Features
abstract
Registration is a fundamental but critical task in point cloud processing, which usually depends on finding element correspondence from two point clouds. However, the finding of reliable correspondence relies on establishing a robust and discriminative description of elements and the correct matching of corresponding elements. In this letter, we develop a coarse-to-fine registration strategy, which utilizes rotation-invariant features in frequency domain and a new graph matching (GM) method for iteratively searching correspondence. In the GM method, the similarity of both nodes and edges in the Euclidean and feature space is formulated to construct the optimization function. The proposed strategy is evaluated using two benchmark datasets and compared with several state-of-the-art methods. Regarding the experimental results, our proposed method can achieve a fine registration with rotation errors of less than 0.2° and translation errors of less than 0.1 m.
Rong Huang 0001, Wei Yao 0008, Yusheng Xu, Zhen Ye 0009, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.5
2021 SOE-Net: A Self-Attention and Orientation Encoding Network for Point Cloud Based Place Recognition
abstract
We tackle the problem of place recognition from point cloud data and introduce a self-attention and orientation encoding network (SOE-Net) that fully explores the relationship between points and incorporates long-range context into point-wise local descriptors. Local information of each point from eight orientations is captured in a PointOE module, whereas long-range feature dependencies among local descriptors are captured with a self-attention unit. Moreover, we propose a novel loss function called Hard Positive Hard Negative quadruplet loss (HPHN quadruplet), that achieves better performance than the commonly used metric learning loss. Experiments on various benchmark datasets demonstrate superior performance of the proposed network over the current state-of-the-art approaches. Our code is released publicly at https://github.com/Yan-Xia/SOE-Net.
Yan Xia 0003, Yusheng Xu, Shuang Li 0008, Rui Wang 0037, Juan Du 0012, Daniel Cremers, Uwe Stilla
CVPR7
2021 Investigation on Misclassification of Pedestrians as Poles by Simulation
abstract
High-precision self-localization is one of the most important capabilities of automated vehicles. Not only accuracy but also localization robustness are crucial for self-driving vehicles in urban environments. The localization robustness decreases by misclassifications of landmarks and therefore false matches between dynamic objects and static landmarks listed in an a priori map. Here we show in the CARLA simulation environment, that the usage of semantic information prevents misclassifications of pedestrians as poles and so increases robustness in urban scenarios. In a simulated scenario of a road intersection pedestrians misclassified without semantic information could be filtered out by class label. In the presented experiments no mismatches of dynamic objects and map landmarks occurred and therefore the localization robustness was increased. Not only pole-like dynamic objects but also semi-static objects like parking cars or freight containers in terminal applications can be detected and excluded from map-based position estimation. The findings of this work show that the introduction of semantic class information leads to a higher self-localization robustness in urban scenarios and therefore should be included into current localization methods.
Christian Rudolf Albrecht, Daniel Névir, Arne-Christoph Hildebrandt, Sven Kraus, Uwe Stilla
IV5
2021 ASFM-Net: Asymmetrical Siamese Feature Matching Network for Point Completion
abstract
We tackle the problem of object completion from point clouds and propose a novel point cloud completion network employing an Asymmetrical Siamese Feature Matching strategy, termed as ASFM-Net. Specifically, the Siamese auto-encoder neural network is adopted to map the partial and complete input point cloud into a shared latent space, which can capture detailed shape prior. Then we design an iterative refinement unit to generate complete shapes with fine-grained details by integrating prior information. Experiments are conducted on the PCN dataset and the Completion3D benchmark, demonstrating the state-of-the-art performance of the proposed ASFM-Net. Our method achieves the 1st place in the leaderboard of Completion3D and outperforms existing methods with a large margin, about 12%. The codes and trained models are released publicly at https://github.com/Yan-Xia/ASFM-Net.
Yaqi Xia, Yan Xia 0003, Wei Li 0111, Rui Song 0003, Kailang Cao, Uwe Stilla
ACM Multimedia6
2021 3-D Point Cloud Generation From Airborne Single-Pass and Single-Channel Circular SAR Data
abstract
This article presents a framework to generate 3-D point clouds by very-high-resolution single-pass and single-channel circular synthetic aperture radar (CSAR). The focus is both on the precise 3-D determination of very small, detached objects and on larger buildings in complex urban scenes. Inspired by optical flow methods, our approach evaluates the tracked aspect dependent backscatter energy flow of objects above the focusing reference plane. Besides the derivation of the exact projection geometry in CSAR, we give answers to the question of how accurate 3-D information can be extracted by this approach as a function of aspect interval and the object’s signal-to-clutter ratio (SCR). A further coherent adaption uses considerably larger subapertures while focusing the data on different reference heights. The computed 3-D information from multiple aspect views is then fused to a georeferenced 3-D point cloud. This results in the first demonstration of 3-D point cloud generation from CSAR data collected with a frequency-modulated continuous-wave (FMCW) radar at the W-band. The results were validated with light detection and ranging (LiDAR) data, and the height accuracy was evaluated in relation to the aspect integration interval and the number of aspect views. For small point targets with edge lengths of 3 cm, we could demonstrate that an aspect interval of only 8.5° leads to a height accuracy below 20 cm. Building roofs reach a height accuracy of several tens of centimeters. Besides the spatial information, we can reveal the angular scattering behavior of individual objects and determine over which aspect interval they are even visible.
Stephan Palm, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.2
2020 Multi-Scale Local Context Embedding for LiDAR Point Cloud Classification
abstract
The semantic interpretation using point clouds, especially regarding light detection and ranging (LiDAR) point cloud classification, has attracted a growing interest in the fields of photogrammetry, remote sensing, and computer vision. In this letter, we aim at tackling a general and typical feature learning problem in 3-D point cloud classification- how to represent geometric features by structurally considering a point and its surroundings in a more effective and discriminative fashion? Recently, enormous efforts have been made to design the geometric features, yet it is less investigated to fully explore the potentials of the features. For that, there have been many filter-based studies proposed by selecting a subset of the whole feature space for better representing the local geometry structure. However, such a hard-threshold selection strategy inevitably suffers from information loss. In addition, the construction of the geometric features is relatively sensitive to the size of the neighborhood. To this end, we propose to extract multi-scaled feature representations and locally embed them into a low-dimensional and robust subspace where a more compact representation with the intrinsic structure preservation of the data is expected to be obtained, thereby further yielding a better classification performance. In our case, we apply a popular manifold learning approach, that is, locality-preserving projections, for the task of learning low-dimensional embedding. Experimental results conducted on one LiDAR point cloud data set provided by the 2018 IEEE Data Fusion Contest demonstrate the effectiveness of the proposed method in comparison with several commonly used state-of-the-art baselines.
Rong Huang 0001, Danfeng Hong, Yusheng Xu, Wei Yao 0008, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.5
2019 Extraction of Multi-Scale Geometric Features for Point Cloud Classification
abstract
Light Detection and Ranging (LiDAR) techniques is an efficient way of obtaining 3D information of complex urban scenes. However, automatically and efficiently interpreting acquired 3D points is still a challenging task. For achieving an excellent semantic interpretation of point clouds, the extraction of distinctive and reliable geometric features often plays a vital role. In this paper, we propose a method generating features from the local vicinity of different sizes and combine them for a better feature representation. To evaluate the proposed method, experiments were conducted using Li-DAR point cloud dataset and compared with that using single scale feature extraction methods.
Rong Huang 0001, Yusheng Xu, Uwe Stilla
IGARSS3
2019 Airborne Circular W-Band SAR for Multiple Aspect Urban Site Monitoring
abstract
This paper presents a strategy for urban site monitoring by very high-resolution circular synthetic aperture radar (CSAR) imaging of multiple aspects. We analytically derive the limits of coherent azimuth processing for nonplanar objects in CSAR if no digital surface model (DSM) is available. The result indicates the level of maximum achievable resolution of these objects in this geometry. The difficulty of constantly illuminating a specific scene in full aspect mode (360°) for such small wavelengths is solved by a hardware- and software-side integration of the radar in a mechanical tracking mode. This results in the first demonstration of full aspect airborne subaperture CSAR images collected with an active frequency-modulated continuous wave (FMCW) radar at W-band. We describe the geometry and the implementation of the real-time beam-steering mode and evaluate resulting effects in the CSAR processing chain. The physical properties in W-band allow the use of extremely short subapertures in length while generating high azimuthal bandwidths. We use this feature to generate full aspect image stacks for CSAR video monitoring in very high frame rates. This technique offers the capability of detecting and observing moving objects in single channel data by shadow tracking. Due to the relatively strong echo of roads, the shadows of moving cars are rich in contrast. The image stack is further evaluated to present wide angular anisotropic properties of targets and first results on multiple aspect image fusion. Both topics show huge potential for further investigations in terms of image analysis and scene classification.
Stephan Palm, Rainer Sommer, Daniel Janssen, Axel Tessmann, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.5
2018 Scale-Awareness of Light Field Camera Based Visual Odometry
Niclas Zeller, Franz Quint, Uwe Stilla
ECCV (8)3
2018 Iterative Calibration of a Vehicle Camera using Traffic Signs Detected by a Convolutional Neural Network
Alexander Hanel, Uwe Stilla
VEHITS2
2018 Voxel-based segmentation of 3D point clouds from construction sites using a probabilistic connectivity model
Yusheng Xu, Sebastian Tuttas, Ludwig Hoegner, Uwe Stilla
Pattern Recognit. Lett.4
2018 Mobile Radar Mapping - Subcentimeter SAR Imaging of Roads
abstract
In this paper, we present a strategy for focusing ultrahigh-resolution synthetic aperture radar (SAR) data for mobile radar mapping. We illustrate the related theoretical background and required extensions on the imaging method based on backprojection techniques. The influence of potential errors in estimating a correct geometry with respect to the imaging quality is investigated in detail by point target simulations. As backprojection techniques require precise knowledge of the topography in close range, the new strategy instantly uses the GPS/INS data of the trajectory to define a suitable digital elevation model of the illuminated scene. We have tested the strategy by driving on conventional roads with an active frequency-modulated continuous wave radar system operating at 300 GHz. Different reference targets were placed in the scene, and the accuracy of the method was evaluated. The results experimentally reveal that the lower terahertz band is capable of subcentimeter SAR imaging in mobile mapping scenarios at very high quality. We have shown that narrow cracks in the asphalt of roads can be detected and fine-scale objects on millimeter size can be displayed. Geometric distortions in the SAR images are significantly reduced allowing the measurements of infrastructure. The output data can at last be transferred to the conventional 3-D Point Cloud Software for further processing.
Stephan Palm, Rainer Sommer, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.3
2017 Semantic segmentation of aerial images with explicit class-boundary modeling
abstract
In this work we propose an end-to-end trainable supervised Deep Convolutional Neural Network (DCNN) targeting the task of semantic-segmentation with the addition of class-aware boundary detection. Through this explicit modeling of the class-boundaries, we enforce the network to extract coherent and complete objects, suppressing the uncertainty influencing these regions. Importantly, we show that class-boundary networks in conjunction with DCNN performs optimally, achieving over 90% overall accuracy (OA) on the challenging ISPRS Vaihingen Semantic Segmentation benchmark.
Dimitrios Marmanis, Konrad Schindler, Jan Dirk Wegner, Mihai Datcu, Uwe Stilla
IGARSS5
2017 Geometric Primitive Extraction From Point Clouds of Construction Sites Using VGS
abstract
We propose a workflow for extracting geometric primitives, including linear, planar, and cylindrical objects, from point clouds of the construction site, using a novel segmentation- and recognition-based strategy. The entire point cloud is first organized by an octree-based voxel structure. The proposed voxel- and graph-based segmentation is conducted by aggregating connected adjacent voxels in a fully connected local affinity graph, the weighted edges of which consider their saliencies simultaneously, including the spatial distance, the shape similarity, and the surface connectivity. After the segmentation, an improved efficient RANSAC algorithm is tailored to recognize and extract geometric primitives from segments. The synthetic, laser scanned, and photogrammetric point clouds are tested in our experiments, and qualitative and quantitative results reveal that our method can outperform the representative segmentation algorithms for our application having the precision and recall better than 0.77. It also shows a good performance with a correctness value better than 0.7 in primitive extraction.
Yusheng Xu, Sebastian Tuttas, Ludwig Hoegner, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.4
2016 Deep Learning Earth Observation Classification Using ImageNet Pretrained Networks
abstract
Deep learning methods such as convolutional neural networks (CNNs) can deliver highly accurate classification results when provided with large enough data sets and respective labels. However, using CNNs along with limited labeled data can be problematic, as this leads to extensive overfitting. In this letter, we propose a novel method by considering a pretrained CNN designed for tackling an entirely different classification problem, namely, the ImageNet challenge, and exploit it to extract an initial set of representations. The derived representations are then transferred into a supervised CNN classifier, along with their class labels, effectively training the system. Through this two-stage framework, we successfully deal with the limited-data problem in an end-to-end processing scheme. Comparative results over the UC Merced Land Use benchmark prove that our method significantly outperforms the previously best stated results, improving the overall accuracy from 83.1% up to 92.4%. Apart from statistical improvements, our method introduces a novel feature fusion algorithm that effectively tackles the large data dimensionality by using a simple and computationally efficient approach.
Dimitrios Marmanis, Mihai Datcu, Thomas Esch, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.4
2015 Establishing a Probabilistic Depth Map from Focused Plenoptic Cameras
abstract
In this paper we propose a novel method for depth estimation based on a single recording of a focused plenoptic camera. The presented algorithm is based on multiple stereo-observations within the multi-view micro images of the focused plenoptic camera. Here, pixel correspondences are found based on local intensity error minimization. Since our algorithm works directly on the micro images, no sub-aperture or epipolar plane images have to be synthesized. Due to the fact that we perform stereo matching based on local criteria we only estimate depth for pixels with sufficient gradient. Thus, we reduce the complexity of the problem, while neglecting uncertain stereo correspondences. Our algorithm incorporates multiple stereo-observations of the same point in a probabilistic depth map. We will show, that this (inverse) depth map can be modeled as a map of Gaussian distributed random variables. Thus, each depth pixel consists of an estimated depth and a corresponding variance, which gives a measure for the uncertainty of the estimation. This uncertainty information can be used in subsequent filtering methods.
Niclas Zeller, Franz Quint, Uwe Stilla
3DV3
2015 An Improved Phase Correlation Method Based on 2-D Plane Fitting and the Maximum Kernel Density Estimator
abstract
In this letter, an improved phase correlation (PC) method based on 2-D plane fitting and the maximum kernel density estimator (MKDE) is proposed, which combines the idea of Stone's method and robust estimator MKDE. The proposed PC method first utilizes a vector filter to minimize the noise errors of the phase angle matrix and then unwraps the filtered phase angle matrix by the use of the minimum cost network flow unwrapping algorithm. Afterward, the unwrapped phase angle matrix is robustly fitted via MKDE, and the slope coefficients of the 2-D plane indicate the subpixel shifts between images. The experiments revealed that the improved method can effectively avoid the impact of outliers on the phase angle matrix during the plane fitting and is robust to aliasing and noise. The matching accuracy can reach 1/50th of a pixel using simulated data. The real image sequence tracking experiment was also undertaken to demonstrate the effectiveness of the proposed PC method with a registration accuracy of root-mean-square error better than 0.1 pixels.
Xiaohua Tong, Yusheng Xu, Zhen Ye 0009, Shijie Liu 0001, Huan Xie 0001, Fengxiang Wang 0002, Sa Gao, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.9
2014 Generating point clouds of forested areas from airborne millimeter wave InSAR data
abstract
Automatic recognition of single trees in remote sensing data is an important research topic in the context of sustainable forest management: In many countries, single-tree related parameters are used as a basis for forest inventory, e.g. tree species, mean tree height or timber volume. Until now, the majority of these parameters are collected manually by measurement of sample plots in cost- and time-intensive field surveys. However, remote sensing-based methods have gained increasing attention during recent decades [1]. In the remote sensing context, an additional application of single tree extraction is driven by the goal to generate information for city tree cadastres or to add layers to geoinformation systems and 3D city models [2].
Michael Schmitt 0003, Uwe Stilla
IGARSS2
2014 Adaptive Multilooking of Airborne Single-Pass Multi-Baseline InSAR Stacks
abstract
Multilooking is a critical task in interferometric synthetic aperture radar (SAR) imaging. While there are many algorithms designed for SAR image pairs and also some first approaches for multi-temporal satellite data stacks, no method suitable to airborne single-pass stacks that typically contain just a small number of multi-baseline acquisitions has been proposed yet. This paper presents an adaptive procedure to determine regions of homogeneous backscattering in heterogeneous scenes such as urban areas. Based on these regions, the complex covariance matrices can be estimated for all pixels in the stack. This step enables the retrieval of all relevant information of the multi-baseline InSAR data set, e.g., despeckled intensity images, interferometric phase observations, and related coherence maps. The denoising efficiency of the proposed method is evaluated and compared to different algorithms. Furthermore, the detail preservation is analyzed in order to prove the validity of the homogeneity assumption.
Michael Schmitt 0003, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.2
2014 Adaptive Covariance Matrix Estimation for Multi-Baseline InSAR Data Stacks
abstract
For many multidimensional applications of synthetic aperture radar (SAR) imaging, the estimation of the covariance matrix for each resolution cell is a critical processing step. The context of this work is the application of covariance matrix estimation for multi-baseline interferometric SAR data sets. In order to ensure local stationarity, which is needed for an unbiased estimation, adaptive techniques are necessary. In this paper, a new approach for adaptive covariance matrix estimation is proposed and evaluated based on measures known from the field of image processing. The procedure is centered around the idea of checking whether the neighboring pixels belong to the same statistical distribution as the currently investigated pixel by applying a threshold to the respective probability density function. All inlier pixels are then used to estimate the complex covariance matrix of the reference pixel. From this covariance matrix, both amplitude and interferometric phase values are extracted, which are then combined for all pixels in the stack in order to employ techniques for the evaluation of filtering efficiency that are typically used in image denoising research. It is found that the proposed algorithm provides high filtering efficiency and good detail preservation at the same time. Apart from that, it is found to be particularly suitable for small-sized stacks of coregistered SAR imagery.
Michael Schmitt 0003, Johannes L. Schönberger, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.3
2013 A sensor-centric EKF for inertial-aided visual odometry
abstract
When appropriate infrastructure is not available, localization of pedestrians becomes a difficult task. This is especially the case in urban or indoor scenarios, where satellite navigation is hindered due to occlusions or multipath effects. A promising alternative is to combine a small, low-cost inertial measurement unit (IMU) with a camera in order to exploit the complementary error characteristics of these devices by simultaneously estimating the positions of observed landmarks and the trajectory of the sensor system with a stochastic filter.
Markus Kleinert, Uwe Stilla
IPIN2
2013 Compressive Sensing Based Layover Separation in Airborne Single-Pass Multi-Baseline InSAR Data
abstract
In this letter, compressive sensing based methods for layover separation in airborne single-pass multi-baseline InSAR data are investigated. The standard compressive sensing (CS) is compared with recently proposed distributed CS (DCS) and a new “multilooking” approach to CS (MCS). Experiments on simulated data show that, while CS is not satisfyingly applicable to single-pass data stacks with just few acquisitions, the neighborhood-based approaches, DCS and MCS, yield a promising perspective.
Michael Schmitt 0003, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.2
2012 Comparison of inertial mechanization approaches for inertial aided monocular EKF-SLAM
Markus Kleinert, Christian Ascher, Uwe Stilla
FUSION3
2012 A concept for building reconstruction from airborne multi-aspect SAR data
abstract
In this work we present a new concept towards building reconstruction based on multi-aspect SAR (MASAR) data where we make use of the inherent redundancy and contradiction of MASAR data. Instead of treating every aspect seperately we perform a joint optimization on all available data. For this purpose we use Probabilistic Graphical Model (PGM) as a declarative representation to encode our knowledge and to reconstruct the most likely objects in the scene.
Oliver Maksymiuk, Uwe Stilla
IGARSS2
2012 First investigations on detection of stationary vehicles in airborne decimeter resolution SAR data by supervised learning
abstract
In this work we investigate the automatic detection of stationary vehicles in SAR images by supervised learning algorithms. This implies the description of the vehicles by a set of representative features. We combine several classes of features including subspace projection based on clustering mechanisms (NMF, PCA), statistical features (image moments), spectral features (gabor wavelets) as well as boundary (shape analysis) and region descriptors (HOG). We further use two different learning algorithms: Support Vector Machines (SVM) and Random Forests.
Oliver Maksymiuk, Michael Schmitt 0003, Andreas R. Brenner, Uwe Stilla
IGARSS4
2012 Adaptive multilooking of airborne Ka-band multi-baseline InSAR data of urban areas
abstract
Multilooking is one of the most important processing steps in SAR interferometry. While formerly fixed-size box-car windows have been used, the problem has become less trivial since decimeter resolution sensors have made a detailed analysis of urban areas possible. This paper presents an approach to detect neighborhoods of homogeneous backscattering in airborne multi-baseline InSAR data stacks. Based on these neighborhoods, an adaptive estimation of the covariance matrix as well as multilooking of interferometric phase and coherence become possible.
Michael Schmitt 0003, Uwe Stilla
IGARSS2
2012 Urban remote sensing by ultra high resolution SAR
abstract
Urban remote sensing by ultra high resolution SAR opens possibilities of recognizing detailed structures and object features beyond the field known from SAR imagery showing meter resolution. SAR data with decimeter spacing or better are a challenge for automatic recognition and reconstruction of detailed structures which were until now extracted from optical data only. This paper addresses current activities on exploiting ultra high resolution SAR images of urban areas.
Uwe Stilla
IGARSS1
2012 Airborne traffic monitoring supported by fast calculated digital surface models
abstract
Vehicle detection in dense urban areas is often complicated due to car-like objects on rooftops which result in false positive detections. This can be easily avoided by using a digital surface model (DSM) calculated from two consecutive images to exclude those regions. However, in the real-time case traffic information has to be gathered rapidly and the calculation of the DSM for the whole image takes a lot of time. The presented approach suggest a method where the disparity image is only calculated for areas of interest. These areas are selected by projecting the road segments from a road database in the original image using the collinearity equation. The local coordinates of the detected vehicles are then transformed back in the UTM coordinate system using the collinearity equation again. It can be shown that the search area for the detector is significantly reduced and which also leads to improved results of the detection.
Sebastian Türmer, Franz Kurz, Peter Reinartz, Uwe Stilla
IGARSS4
2012 On sensor pose parameterization for inertial aided visual SLAM
abstract
When appropriate infrastructure is not available, localization of pedestrians becomes a difficult task. This is especially the case in urban or indoor scenarios, where satellite navigation is hindered due to occlusions or multipath effects. A promising alternative is to combine a small, low-cost IMU with a camera in order to exploit the complementary error characteristics of these devices by simultaneously estimating the positions of observed landmarks and the trajectory of the sensor system with a stochastic filter. In this work, a standard approach to parameterize the error in position and attitude estimates that is commonly used in GNSS-INS integration is compared to alternative parameterizations that are based on the twist representation of rigid body motions, which has gained increasing popularity in the literature. For this purpose, the error-state transition and measurement equations are formulated for the twist representation as well as for the standard approach. Finally, the different approaches are compared on a simulated and a real indoor dataset by applying an extended Kalman filter (EKF).
Markus Kleinert, Uwe Stilla
IPIN2
2012 Simultaneous Calibration of ALS Systems and Alignment of Multiview LiDAR Scans of Urban Areas
abstract
Tasks such as city modeling or urban planning require the registration, alignment, and comparison of multiview and/or multitemporal remote sensing data. Airborne laser scanning (ALS) is one of the established techniques to deliver these data. Regrettably, direct georeferencing of ALS measurements usually leads to considerable displacements that limit connectivity and/or comparability of overlapping point clouds. Most reasons for this effect can be found in the impreciseness of the positioning and orientation sensors and their misalignment to the laser scanner. Typically, these sensors are comprised of a global navigation satellite system receiver and an inertial measurement unit. This paper presents a method for the automatic self-calibration of such ALS systems and the alignment of the acquired laser point clouds. Although applicable to classical nadir configurations, a novelty of our approach is the consideration of multiple data sets that were recorded with an oblique forward-looking full-waveform laser scanner. A combination of a region-growing approach with a random-sample-consensus segmentation method is used to extract planar shapes. Matching objects in overlapping data sets are identified with regard to several geometric attributes. A new methodology is presented to transfer the planarity constraints into systems of linear equations to determine both the boresight parameters and the data alignment. In addition to system calibration and data registration, the presented workflow results in merged 3-D point clouds that contain information concerning rooftops and all building facades. This database represents a solid basis and reference for applications such as change detection.
Marcus Hebel, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.2
2011 Comparison of Two Methods for Vehicle Extraction From Airborne LiDAR Data Toward Motion Analysis
abstract
It has been revealed that single-pass airborne light detection and ranging (LiDAR) system (ALS) data could provide not only the spatial but also the dynamical information of a scanned scene due to the so-called motion artifact effect. A common strategy for extracting dynamical information from ALS data is established based on analyzing shape deformations of vehicles which have to be extracted in advance. Therefore, vehicle extraction results are directly related to the performance of motion analysis. In this letter, two vehicle extraction methods, namely, grid-cell- and 3-D point-cloud-analysis-based methods, which represent two main streams in LiDAR data processing, are to be evaluated and compared toward influences on the performance of motion analysis. Motion estimation based on the two methods is respectively applied to real ALS data sets. The results show that the 3-D data-based method can yield more accurate and robust dynamical traffic information such as motion state and velocity of vehicles, while the grid-cell-based method can provide more complete information by extracting more stationary vehicles.
Wei Yao 0008, Uwe Stilla
IEEE Geosci. Remote. Sens. Lett.2
2010 The Effects of Radiometry on the Accuracy of Intensity Based Registration
abstract
Besides several other factors, radiometric differences between a reference and a floating image greatly influence the achievable accuracy of image registration. In this work we derive the magnitude of registration inaccuracy coming from changes in radiometric properties. This is done for the example of medical X-ray image registration. We therefore estimate the change of image intensity with respect to object shape, X-ray attenuation of the object material and the initial X-ray energy by modeling a simplified image formation process. The change in intensity is then used to determine a closed form estimation of the resulting registration error, independent from a specific registration algorithm. Finally the theoretical calculations are compared to the accuracy of intensity based registration performed on X-ray images with different radiometric properties. Results show that the herewith derived accuracy estimation is well suited to predict the achievable accuracy of a registration for images with radiometric differences.
Boris Peter Selby, Georgios Sakas, Stefan Walter, Wolf-Dieter Groch, Uwe Stilla
ICPR5
2010 Utilization of airborne multi-aspect InSAR data for the generation of urban ortho-images
abstract
This paper addresses the generation of “true” RADAR ortho-images from highest resolution multi-aspect InSAR data. Due to the side-looking SAR imaging geometry, the well-known layover and shadowing effects prevent the production of truly rectified ortho-imagery from one image alone. Here, an approach for the reconstruction of Digital Surface Models of densely built inner city areas is proposed. Since in SAR interferometry each dataset pixel contains not only the interferometric phase needed for 3D reconstruction but also the corresponding amplitude or intensity value, respectively, the procedure can also be seen in the context of true ortho-rectification.
Michael Schmitt 0003, Uwe Stilla
IGARSS2
2010 Extraction of building polygons from SAR images: Grouping and decision-level in the GESTALT system
Eckart Michaelsen, Uwe Stilla, Uwe Sörgel, Leo J. Doktorski
Pattern Recognit. Lett.2
2010 Automatic vehicle extraction from airborne LiDAR data of urban areas aided by geodesic morphology
Wei Yao 0008, Stefan Hinz, Uwe Stilla
Pattern Recognit. Lett.3
2010 Road Network Extraction in VHR SAR Images of Urban and Suburban Areas by Means of Class-Aided Feature-Level Fusion
abstract
In this paper, we propose to combine two road extractors from very high resolution synthetic aperture radar scenes: one more successful in rural areas and one explicitly designed for urban areas. In order to get the best combination of both, a rapid mapping filter for discriminating rural and urban scenes is utilized. Finally, the results are fused on a feature level and connected by means of a network optimization. The approach is tested and evaluated on TerraSAR-X data containing complex urban areas and urban-rural fringe scenes.
Karin Hedman, Uwe Stilla, Gianni Lisini, Paolo Gamba
IEEE Trans. Geosci. Remote. Sens.2
2010 Vehicle Detection in Very High Resolution Satellite Images of City Areas
abstract
Current traffic research is mostly based on data from fixed-installed sensors like induction loops, bridge sensors, and cameras. Thereby, the traffic flow on main roads can partially be acquired, while data from the major part of the entire road network are not available. Today's optical sensor systems on satellites provide large-area images with 1-m resolution and better, which can deliver complement information to traditional acquired data. In this paper, we present an approach for automatic vehicle detection from optical satellite images. Therefore, hypotheses for single vehicles are generated using adaptive boosting in combination with Haar-like features. Additionally, vehicle queues are detected using a line extraction technique since grouped vehicles are merged to either dark or bright ribbons. Utilizing robust parameter estimation, single vehicles are determined within those vehicle queues. The combination of implicit modeling and the use ofa prioriknowledge of typical vehicle constellation leads to an enhanced overall completeness compared to approaches which are only based on statistical classification techniques. Thus, a detection rate of over 80% is possible with very high reliability. Furthermore, an approach for movement estimation of the detected vehicle is described, which allows the distinction of moving and stationary traffic. Thus, even an estimate for vehicles' speed is possible, which gives additional information about the traffic condition at image acquisition time.
Jens Leitloff, Stefan Hinz, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.3
2010 Mutual Enhancement of Weak Laser Pulses for Point Cloud Enrichment Based on Full-Waveform Analysis
abstract
In this paper, we present a novel method for weak laser pulse detection by full-waveform analysis. Pulse detection is a fundamental step of processing data of pulsed laser systems for extracting features of the illuminated object. Weak laser pulses below the threshold are discarded by classical methods. For full-waveform laser scanners, the entire recording of a scene can be interpreted as a discrete waveform cuboid I[x y t]-, where the measured amplitude at each time t and each beam direction [x y] is stored. The potential information hiding in the waveform cuboid could be utilized to improve the analysis result of conventional system. The neighborhood relation given by co-planarity constraint in waveform data is analyzed. Waveform stacking technique is introduced to improve the signalto-noise ratio (SNR) of objects with poor surface response in view of mutual information enhancement. Hypotheses for planar surface of different slopes are generated and verified. Each pulse signal is assessed with respect to accepted hypotheses by a contribution measure to the local geometry. Pulse signals are finally classified according to the likelihood value by automatically thresholding. The presented method was applied to waveform data of an urban scene and showed very promising results. The pulses reflected from objects with poor surface response or partially occluded are redetected, which cannot be predicted by given geometric models based on available points.
Wei Yao 0008, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.2
2009 Hybrid GPU-Based Single- and Double-Bounce SAR Simulation
abstract
In this paper, a new hybrid graphics-processing-unit (GPU)-based real-time synthetic aperture radar (SAR) simulation system is presented. Previous real-time SAR simulators only supported single-bounce simulation in real time. The new hybrid system uses the rasterization approach for real-time single-bounce simulation and a new image-based GPU ray-tracing approach for monostatic SAR double-bounce simulation. This approach provides fast simulation results even while simulating complex and extended scenes. The simulation results are compared to a high-resolution airborne SAR image, and the limitations of the approach are discussed.
Timo Balz, Uwe Stilla
IEEE Trans. Geosci. Remote. Sens.2
2008 Internal evaluation of registration results for radiographic images
abstract
This wort focuses on internal gray level based evaluation of image registration results. The motivation is to provide all approach for self-diagnosis in the scope of a patient alignment system based on rigid registration of real and reconstructed X-ray images. As an automatic system should provide expressive indicators for the correctness of the outcome, we propose a method to estimate the probability for the resulting transformations to lie within a predefined window of acceptable values. Based purely on image gray values, the approach is independent from previous knowledge about the images. By registration of corresponding fragments of both images we generate redundancy and define the probability density of the resulting transformations. The proposed method is tested comparing digital reconstructed radiographs (DRRs) to X-ray images. By introducing geometric and radiometric deviations we show that a reliable self-diagnosis is possible.
Boris Peter Selby, Georgios Sakas, Stefan Walter, Wolf-Dieter Groch, Uwe Stilla
ICPR5
2008 Evaluation of a Statistical Fusion of Linear Features in SAR Data
abstract
In this paper, we describe an extension of an automatic road extraction procedure developed for single SAR images towards multi-aspect SAR images. Extracted information from multi-aspect SAR images is not only redundant and complementary, in some cases even contradictory. Hence, multi-aspect SAR images require a careful selection within the fusion step. In this work, a fusion step based on probability theory is proposed. During fusion each extracted line primitive is assessed by means of Bayesian probability theory. The assessment is based on the attributes of the line primitive (i.e. length, straightness, etc), global context and sensor geometry. The fusion and its integration into the road extraction system are tested in a sub-urban SAR scene.
Karin Hedman, Stefan Hinz, Uwe Stilla
IGARSS (4)3
2007 Building feature extraction via a deterministic approach: application to real high resolution SAR images
abstract
Interpretation of high resolution SAR (synthetic aperture radar) images is still a hard task, especially when man-made objects crowd the scene under detection. This paper contributes to the analysis of this kind of data by adopting an approach, based on a scattering model, for the retrieval of buildings height from real SAR images and presenting first numerical results.
Giorgio Franceschetti, Raffaella Guida, Antonio Iodice, Daniele Riccio, Giuseppe Ruello, Uwe Stilla
IGARSS6
2006 Exploiting Multi-Aspect SAR Data for Object Extraction
abstract
In this paper we describe a fusion approach for automatic object extraction from multi-aspect SAR images. Before fusion the uncertainty of each extracted object is assessed by means of Bayesian probability theory. The assessment is performed on attribute-level and is based on predefined probability density functions learned from training data. I. INTRODUCTION Automatic extraction of man-made objects from synthetic aperture radar (SAR) images is regarded as a complicated task. Compared to optical image acquisition, SAR system is an active system and can operate during day and night. It is also nearly weather-independent and, moreover, during bad weather conditions, SAR is the only operational system available today. Extraction of man-made objects from SAR images therefore offers a suitable complement or alternative to object extraction from optical images. The recent development of new high resolution SAR systems offers new potential for automatic object extraction. Satellite SAR images up to 1 m resolution will soon be available by the launch of the German satellite TerraSAR-X (1). Airborne images already provide resolution up to 1 decimetre (2). However, the improved resolution does not automatically make automatic object extraction easier, yet it faces new challenges. Especially in urban areas, the complexity arises through dominant scattering caused by building structures, traffic signs and metallic objects in cities. These bright features hinder important extractable features. The inevitable consequences of the side-looking geometry of SAR, occlusions caused by shadow- and layover effects, is present in forestry areas as well as in built-up areas. In urban areas, the best results for the visibility of roads are obtained, when the illumination direction coincide with the main road orientations (3). Preliminary work has shown that the usage of SAR images illuminated from different directions (i.e. multi- aspect images) improves the road extraction results. This has been tested both for real and simulated SAR scenes (4)(5). Multi-aspect SAR images has appeared to be an interesting topic for automatic building extraction as well (6). In this article we present a fusion concept for object extraction based on a Bayesian statistical approach, which incorporates both global context and sensor geometry. The fusion will be implemented in a road extraction approach, (Sect. II), but can as well be applied for other man-made objects. The main focus of this paper is the proposed fusion module, which is explained in Sect. III. Some intermediate results of an uncertainty assessment of line segments based on a training step and global context are discussed in Sect IV. II. ROAD EXTRACTION SYSTEM The extraction of roads from SAR images is based on an already existing road extraction approach (7), which was originally designed for optical images with a ground pixel size of about 2m (8). The first step consists of line extraction using Steger's differential geometry approach (9), which is followed by a smoothening and splitting step. By applying explicit knowledge about roads, the line segments are evaluated according to their attributes such as width, length, curvature, etc. The evaluation is performed within the fuzzy theory. A weighted graph of the evaluated road segments is constructed. For the extraction of the roads from the graph, supplementary road segments are introduced and seed points are defined. Best- valued road segments serve as seed points, which are connected by an optimal path search through the graph. The novelty presented in this paper refers on one hand to the adoption of the fusion module to multi-aspect SAR images and on the other hand to a probabilistic formulation of the fusion problem instead of using fuzzy-functions.
Karin Hedman, Stefan Hinz, Uwe Stilla
IGARSS3
2006 Car detection in aerial thermal images by local and global evidence accumulation
Stefan Hinz, Uwe Stilla
Pattern Recognit. Lett.2
2005 Context-supported vehicle detection in optical satellite images of urban areas
abstract
ABSTRACT: Due to increasing traffic there is high demand in traffic monitoring of densely populated urban areas. In our approach we focus on the detection of vehicle queues and use a priori information of roads location and direction. In high resolution satellite imagery single vehicles can hardly be separated since they are merged to either dark or bright ribbons. Initial hypotheses for the queues can be extracted as lines in scale space which represent the centres of the queues. We exploit the context information that vehicle queues are composed of repetitive patterns. For discrimination of single vehicles a width function of the queues is calculated in the gradient image and the variations of the width function are analyzed. We show intermediate and final results of processing a panchromatic QuickBird image covering a part of an inner city area. ‡ A preliminary version of this article has been presented at the ISPRS-Workshop on “High Resolution Earth Imaging for Geo-Information”, Hanover, 2005.
Stefan Hinz, Jens Leitloff, Uwe Stilla
IGARSS3
2003 Elimination of across-track phase components in airborne along-track interferometry data to improve object velocity measurements
abstract
The phase differences in airborne along-track interferometry can be used to calculate radial velocities of moving objects. This paper deals with the problem to preprocess phase and intensity data of along-track interferometry with the aim to determine velocities of slow moving objects. The phase information is often severely disturbed by quite different reasons. Beside the problem of noise sensitivity of interferometric measurements in airborne measurements additionally an unwanted across-track component could lead to systematic phase errors. We propose a first order correction of the along-track phase values by subtracting a plane which is fitted into phase values. In case of a flat scene this correction complies a 'flat Earth' correction of the phase. The described process was applied to a scene with slow moving cargo ships.
Karsten Schulz, Uwe Sörgel, Ulrich Thoennessen, Uwe Stilla
IGARSS4
2003 Determination of optimal SAR illumination aspects in build-up areas
abstract
The increasing resolution of SAR data opens the possibility to utilize such data for scene interpretation in urban areas. SAR specific illumination phenomena like foreshortening, layover, shadow, and multipath-propagation burden the interpretation. In this paper a high resolution LIDAR DEM is incorporated to determine the visibility of objects by a SAR measurement from a given sensor trajectory and orientation in an urban environment. Shadow and layover areas are detected by incoherent sampling of the DEM. By a variation of viewing angle and aspect a large number of such simulations are carried out. From this set a subset of a number of best simulations is determined for to two example tasks, namely the analysis of objects on roads and the detection and reconstruction of buildings. Dominant scattering can interfere large parts of the scene. Such strong backscatter can be caused by specular reflection on gabled roofs or double-bounce on corner structures at building walls and roofs. The approach is extended to the detection of such structures.
Uwe Sörgel, Karsten Schulz, Ulrich Thoennessen, Uwe Stilla
IGARSS4
2003 Perceptual grouping of regular structures for automatic detection of man-made objects
abstract
Human observers perceive man-made objects in images from the visual spectrum domain as well as in IR or SAR imagery. Mechanisms like perceptual grouping are crucial to this capability. In this paper two examples for grouping in different image sources are discussed. The first example is activity estimation in urban areas from thermal IR images. The grouping of vehicles into rows is performed along the margins of the roads. The other example is related to the detection of industrial buildings from InSAR data. Such buildings often show salient regular patterns of strong scatterers on their roofs. A previous segmentation which uses the intensity, height and coherence information extracts building cues. Strong scatterers are filtered by a spot detector and localized by a cluster formation. These scatterers are grouped in rows by a process that uses the contours of the building cues as context.
Uwe Stilla, Eckart Michaelsen, Uwe Sörgel, Karsten Schulz
IGARSS1