Uwe Franke

dblp:56/2085 · DBLP profile ↗
← Back
52ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 53% Segmentation and scene understanding · 40% Autonomous driving · 5%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.042019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
Tree-Structured Models for Efficient Multi-Cue Scene Labeling · IEEE Trans. Pattern Anal. Mach. Intell. 2017
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › 3D vision
stereo vision
0.942019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice · ECCV (2) 2014
Computer vision › Segmentation and scene understanding
instance segmentation
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › 3D vision
stereo video dataset
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Robotics › Autonomous driving
urban scene understanding
0.212016
The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016
Computer vision › 3D vision
motion estimation
0.222011
Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011
Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time · ECCV (4) 2010
Computer vision › 3D vision
scene flow estimation
0.222011
Stereoscopic Scene Flow Computation for 3D Motion Understanding · Int. J. Comput. Vis. 2011
Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008
Computer vision › 3D vision
depth estimation
0.212014
Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice · ECCV (2) 2014
Computer vision › Segmentation and scene understanding › semantic segmentation › efficient semantic segmentation
real-time semantic segmentation
0.212014
Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding · ECCV (5) 2014
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding
0.212014
Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding · ECCV (5) 2014
Computer vision › 3D vision › stereo vision
stereo matching
0.222008
Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008
Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007
Computer vision › 3D vision › motion estimation
optical flow
0.112011
Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011
Computer vision › 3D vision › motion estimation › optical flow
variational optical flow
0.112011
Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011
Computer vision › Segmentation and scene understanding
scene understanding
0.112019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
Computer vision › 3D vision › stereo vision
stereo image sequences
0.112010
Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time · ECCV (4) 2010
Computer vision › 3D vision › scene flow estimation
dense scene flow
0.112008
Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008
Computer vision › 3D vision
3d reconstruction
0.112007
Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007
Computer vision › 3D vision › stereo vision › stereo matching
dense stereo matching
0.112007
Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007
Computer vision › 3D vision › stereo vision › stereo matching › disparity refinement
sub-pixel disparity estimation
0.112007
Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007
Computer vision › 3D vision › motion estimation
3d motion estimation
0.012011
Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011
Computer vision › 3D vision › stereo vision › stereo matching
semi-global matching
0.012007
Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007
Image and video coding › transform coding
adaptive transform coding
0.011992
Spectral Entropy-Activity Classification in Adaptive Transform Coding · IEEE J. Sel. Areas Commun. 1992
Image and video coding
transform coding
0.011992
Spectral Entropy-Activity Classification in Adaptive Transform Coding · IEEE J. Sel. Areas Commun. 1992

Methods — techniques the papers use, named apart from their topics

over-segmentation · 0.4global energy minimization · 0.4fully convolutional network · 0.4tree-structured model · 0.3superpixel · 0.3conditional random field · 0.3weak annotation · 0.2deep learning · 0.2real-time inference · 0.2medium-level model · 0.2spectral entropy · 0.0activity measurement · 0.0
YearPublicationVenuePosition
2021 Learning Stixel-based Instance Segmentation
abstract
Stixels have been successfully applied to a wide range of vision tasks in autonomous driving, recently including instance segmentation. However, due to their sparse occurrence in the image, until now Stixels seldomly served as input for Deep Learning algorithms, restricting their utility for such approaches. In this work we present StixelPointNet, a novel method to perform fast instance segmentation directly on Stixels. By regarding the Stixel representation as unstructured data similar to point clouds, architectures like PointNet are able to learn features from Stixels. We use a bounding box detector to propose candidate instances, for which the relevant Stixels are extracted from the input image. On these Stixels, a PointNet models learns binary segmentations, which we then unify throughout the whole image in a final selection step. StixelPointNet achieves state-of-the-art performance on Stixel-level, is considerably faster than pixel-based segmentation methods, and shows that with our approach the Stixel domain can be introduced to many new 3D Deep Learning tasks.
Monty Santarossa, Lukas Schneider, Claudius Zelenka, Lars Schmarje, Reinhard Koch, Uwe Franke
IV6
2020 Single-Shot 3D Detection of Vehicles from Monocular RGB Images via Geometrically Constrained Keypoints in Real-Time
abstract
In this paper we propose a novel 3D single-shot object detection method for detecting vehicles in monocular RGB images. Our approach lifts 2D detections to 3D space by predicting additional regression and classification parameters and hence keeping the runtime close to pure 2D object detection. The additional parameters are transformed to 3D bounding box keypoints within the network under geometric constraints. Our proposed method features a full 3D description including all three angles of rotation without supervision by any labeled ground truth data for the object's orientation, as it focuses on certain keypoints within the image plane. While our approach can be combined with any modern object detection framework with only little computational overhead, we exemplify the extension of SSD for the prediction of 3D bounding boxes. We test our approach on different datasets for autonomous driving and evaluate it using the challenging KITTI 3D Object Detection as well as the novel nuScenes Object Detection benchmarks. While we achieve competitive results on both benchmarks we outperform current state-of-the-art methods in terms of speed with more than 20 FPS for all tested datasets and image resolutions.
Nils Gählert, Jun-Jun Wan, Nicolas Jourdan 0001, Jan Finkbeiner, Uwe Franke, Joachim Denzler
IV5
2019 Beyond Bounding Boxes: Using Bounding Shapes for Real-Time 3D Vehicle Detection from Monocular RGB Images
abstract
The representation of objects as 2D bounding boxes in monocular RGB images limits the faculty of current computer vision systems to 2D object detection. It fails to provide crucial information such as the orientation of other vehicles, which is vital for autonomous driving. At the same time, real-time performance is essential to qualify an approach for deployment in a productive environment. In order to tackle this problem, we present an approach that predicts several key points selected from a virtual 3D bounding box around a vehicle instead of a pure 2D bounding box. These key points can be interpreted as a bounding shape. With this novel representation we can calculate the actual 3D bounding box of the corresponding object. Thanks to the straightforward implementation of bounding shape in any current state-of-the-art 2D object detector both for singleshot frameworks like YOLO or SSD as well as for two-stage detectors like Faster-RCNN with a minimum of computational overhead, it is able to be run in real-time while providing additional useful information for vehicle detection. We exemplify the extension of SSD to Bounding Shape SSD ( BS3D) and evaluate our approach using the challenging KITTI as well as the novel VIPER dataset.
Nils Gählert, Jun-Jun Wan, Michael Weber 0009, Johann Marius Zöllner, Uwe Franke, Joachim Denzler
IV5
2019 Slanted Stixels: A Way to Represent Steep Streets
abstract
Abstract This work presents and evaluates a novel compact scene representation based on Stixels that infers geometric and semantic information. Our approach overcomes the previous rather restrictive geometric assumptions for Stixels by introducing a novel depth model to account for non-flat roads and slanted objects. Both semantic and depth cues are used jointly to infer the scene representation in a sound global energy minimization formulation. Furthermore, a novel approximation scheme is introduced in order to significantly reduce the computational complexity of the Stixel algorithm, and then achieve real-time computation capabilities. The idea is to first perform an over-segmentation of the image, discarding the unlikely Stixel cuts, and apply the algorithm only on the remaining Stixel cuts. This work presents a novel over-segmentation strategy based on a fully convolutional network, which outperforms an approach based on using local extrema of the disparity map. We evaluate the proposed methods in terms of semantic and geometric accuracy as well as run-time on four publicly available benchmark datasets. Our approach maintains accuracy on flat road scene datasets while improving substantially on a novel non-flat road dataset.
Daniel Hernández Juárez, Lukas Schneider, Pau Cebrian, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys, Juan C. Moure
Int. J. Comput. Vis.7
2018 MB-Net: MergeBoxes for Real-Time 3D Vehicles Detection
abstract
High performance vehicle detection and pose esti- mation in RGB images is essential for driver assistance systems as well as for autonomous vehicles. Classical 2D box-based detection schemes allow roughly estimating the position of other vehicles, but not their orientation relative to the ego-vehicle. Recent approaches use 3D models to derive the pose of other vehicles from single monocular images but do not reach real- time performance. In this paper we present an approach that achieves competitive performance on the challenging KITTI Object Detection and orientation Estimation benchmark while being the fastest approach with over 40 FPS. The key is a novel representation named MergeBox whose parameters can be estimated extremely efficiently. We extend SSD-a current fast state-of-the-art 2D box object detector- with this representation to our MB-Net. In contrast to all other current state-of-the-art methods we do not require explicit information on the object orientation for training our model. This reduces label costs significantly, a further advantage for practical applications that require labeling of databases that are much bigger than those used for research.
Nils Gählert, Marina Mayer, Lukas Schneider, Uwe Franke, Joachim Denzler
Intelligent Vehicles Symposium4
2018 Box2Pix: Single-Shot Instance Segmentation by Assigning Pixels to Object Boxes
abstract
The task of semantic instance segmentation has gained a large interest within academia as well as industry, especially in the context of autonomous driving. While several published approaches achieve very strong results, only few of them achieve frame rates that are sufficient for the automotive domain. We present an approach that achieves competitive results on the Cityscapes [1] and KITTI [2] datasets, while being twice as fast as any other existing approach. Our method relies on a single fully-convolutional network (FCN [3]) predicting object bounding boxes, as well as pixel-wise semantic object classes and an offset vector pointing to corresponding object centers. Using those outputs, we present an efficient and simple post-processing that assigns each object pixel to its best matching object detection, resulting in an instance segmentation obtained at real-time speeds.
Jonas Uhrig, Eike Rehder, Björn Fröhlich, Uwe Franke, Thomas Brox
Intelligent Vehicles Symposium4
2017 Sparsity Invariant CNNs
abstract
In this paper, we consider convolutional neural networks operating on sparse inputs with an application to depth completion from sparse laser scan data. First, we show that traditional convolutional networks perform poorly when applied to sparse data even when the location of missing data is provided to the network. To overcome this problem, we propose a simple yet effective sparse convolution layer which explicitly considers the location of missing data during the convolution operation. We demonstrate the benefits of the proposed network architecture in synthetic and real experiments with respect to various baseline approaches. Compared to dense baselines, the proposed sparse convolution network generalizes well to novel datasets and is invariant to the level of sparsity in the data. For our evaluation, we derive a novel dataset from the KITTI benchmark, comprising over 94k depth annotated RGB images. Our dataset allows for training and evaluating depth completion and depth prediction techniques in challenging real-world settings and is available online at: www.cvlibs.net/datasets/kitti.
Jonas Uhrig, Nick Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, Andreas Geiger 0001
3DV4
2017 Slanted Stixels: Representing San Francisco's Steepest Streets
Daniel Hernández Juárez, Lukas Schneider, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys
BMVC7
2017 Detecting unexpected obstacles for self-driving cars: Fusing deep learning and geometric modeling
abstract
The detection of small road hazards, such as lost cargo, is a vital capability for self-driving cars. We tackle this challenging and rarely addressed problem with a vision system that leverages appearance, contextual as well as geometric cues. To utilize the appearance and contextual cues, we propose a new deep learning-based obstacle detection framework. Here a variant of a fully convolutional network is proposed to predict a pixel-wise semantic labeling of (i) free-space, (ii) on-road unexpected obstacles, and (iii) background. The geometric cues are exploited using a state-of-the-art detection approach that predicts obstacles from stereo input images via model-based statistical hypothesis tests. We present a principled Bayesian framework to fuse the semantic and stereo-based detection results. The mid-level Stixel representation is used to describe obstacles in a flexible, compact and robust manner. We evaluate our new obstacle detection system on the Lost and Found dataset, which includes very challenging scenes with obstacles of only 5 cm height. Overall, we report a major improvement over the state-of-the-art, with a performance gain of 27.4%. In particular, we achieve a detection rate of over 90% for distances of up to 50 m. Our system operates at 22 Hz on our self-driving platform.
Sebastian Ramos, Stefan K. Gehrig, Peter Pinggera, Uwe Franke, Carsten Rother
Intelligent Vehicles Symposium4
2017 RegNet: Multimodal sensor registration using deep neural networks
abstract
In this paper, we present RegNet, the first deep convolutional neural network (CNN) to infer a 6 degrees of freedom (DOF) extrinsic calibration between multimodal sensors, exemplified using a scanning LiDAR and a monocular camera. Compared to existing approaches, RegNet casts all three conventional calibration steps (feature extraction, feature matching and global regression) into a single real-time capable CNN. Our method does not require any human interaction and bridges the gap between classical offline and target-less online calibration approaches as it provides both a stable initial estimation as well as a continuous online correction of the extrinsic parameters. During training we randomly decalibrate our system in order to train RegNet to infer the correspondence between projected depth measurements and RGB image and finally regress the extrinsic calibration. Additionally, with an iterative execution of multiple CNNs, that are trained on different magnitudes of decalibration, our approach compares favorably to state-of-the-art methods in terms of a mean calibration error of 0.28° for the rotational and 6 cm for the translation components even for large decalibrations up to 1.5 m and 20°.
Nick Schneider, Florian Piewak, Christoph Stiller, Uwe Franke
Intelligent Vehicles Symposium4
2017 The Stixel World: A medium-level representation of traffic scenes
Marius Cordts, Timo Rehfeld, Lukas Schneider, David Pfeiffer, Markus Enzweiler, Stefan Roth 0001, Marc Pollefeys, Uwe Franke
Image Vis. Comput.8
2017 Stereo vision during adverse weather - Using priors to increase robustness in real-time stereo vision
Stefan K. Gehrig, Nicolai Schneider, Reto Stalder, Uwe Franke
Image Vis. Comput.4
2017 Tree-Structured Models for Efficient Multi-Cue Scene Labeling
abstract
We propose a novel approach to semantic scene labeling in urban scenarios, which aims to combine excellent recognition performance with highest levels of computational efficiency. To that end, we exploit efficient tree-structured models on two levels: pixels and superpixels. At the pixel level, we propose to unify pixel labeling and the extraction of semantic texton features within a single architecture, so-called encode-and-classify trees. At the superpixel level, we put forward a multi-cue segmentation tree that groups superpixels at multiple granularities. Through learning, the segmentation tree effectively exploits and aggregates a wide range of complementary information present in the data. A tree-structured CRF is then used to jointly infer the labels of all regions across the tree. Finally, we introduce a novel object-centric evaluation method that specifically addresses the urban setting with its strongly varying object scales. Our experiments demonstrate competitive labeling performance compared to the state of the art, while achieving near real-time frame rates of up to 20 fps.
Marius Cordts, Timo Rehfeld, Markus Enzweiler, Uwe Franke, Stefan Roth 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 The Cityscapes Dataset for Semantic Urban Scene Understanding
abstract
Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately captures the complexity of real-world urban scenes. To address this, we introduce Cityscapes, a benchmark suite and large-scale dataset to train and test approaches for pixel-level and instance-level semantic labeling. Cityscapes is comprised of a large, diverse set of stereo video sequences recorded in streets from 50 different cities. 5000 of these images have high quality pixel-level annotations, 20 000 additional images have coarse annotations to enable methods that leverage large volumes of weakly-labeled data. Crucially, our effort exceeds previous attempts in terms of dataset size, annotation richness, scene variability, and complexity. Our accompanying empirical study provides an in-depth analysis of the dataset characteristics, as well as a performance evaluation of several state-of-the-art approaches based on our benchmark.
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth 0001, Bernt Schiele
CVPR7
2016 Visual odometry driven online calibration for monocular LiDAR-camera systems
abstract
Recently LiDAR-camera systems have rapidly emerged in many applications. The integration of laser range-finding technologies into existing vision systems enables a more comprehensive understanding of 3D structure of the environment. The advantage, however, relies on a good geometrical calibration between the LiDAR and the image sensors. In this paper we consider visual odometry, a discipline in computer vision and robotics, in the context of recently emerging online sensory calibration studies. By embedding the online calibration problem into a LiDAR-monocular visual odometry technique, the temporal change of extrinsic parameters can be tracked and compensated effectively.
Hsiang-Jen Chien, Reinhard Klette, Nick Schneider, Uwe Franke
ICPR4
2016 Lost and Found: detecting small road hazards for self-driving vehicles
abstract
Detecting small obstacles on the road ahead is a critical part of the driving task which has to be mastered by fully autonomous cars. In this paper, we present a method based on stereo vision to reliably detect such obstacles from a moving vehicle.
Peter Pinggera, Sebastian Ramos, Stefan K. Gehrig, Uwe Franke, Carsten Rother, Rudolf Mester
IROS4
2016 Semantic Stixels: Depth is not enough
abstract
In this paper we present Semantic Stixels, a novel vision-based scene model geared towards automated driving. Our model jointly infers the geometric and semantic layout of a scene and provides a compact yet rich abstraction of both cues using Stixels as primitive elements. Geometric information is incorporated into our model in terms of pixel-level disparity maps derived from stereo vision. For semantics, we leverage a modern deep learning-based scene labeling approach that provides an object class label for each pixel. Our experiments involve an in-depth analysis and a comprehensive assessment of the constituent parts of our approach using three public benchmark datasets. We evaluate the geometric and semantic accuracy of our model and analyze the underlying run-times and the complexity of the obtained representation. Our results indicate that the joint treatment of both cues on the Semantic Stixel level yields a highly compact environment representation while maintaining an accuracy comparable to the two individual pixel-level input data sources. Moreover, our framework compares favorably to related approaches in terms of computational costs and operates in real-time.
Lukas Schneider, Marius Cordts, Timo Rehfeld, David Pfeiffer, Markus Enzweiler, Uwe Franke, Marc Pollefeys, Stefan Roth 0001
Intelligent Vehicles Symposium6
2016 Context-based multi-target tracking with occlusion handling
Junli Tao, Uwe Franke, Reinhard Klette
Mach. Vis. Appl.2
2015 What Is in Front? Multiple-Object Detection and Tracking with Dynamic Occlusion Handling
Junli Tao, Markus Enzweiler, Uwe Franke, David Pfeiffer, Reinhard Klette
CAIP (1)3
2015 High-performance long range obstacle detection using stereo vision
abstract
Reliable detection of obstacles at long range is crucial for the timely response to hazards by fast-moving safety-critical platforms like autonomous cars. We present a novel method for the joint detection and localization of distant obstacles using a stereo vision system on a moving platform. The approach is applicable to both static and moving obstacles and pushes the limits of detection performance as well as localization accuracy.
Peter Pinggera, Uwe Franke, Rudolf Mester
IROS2
2015 Low-level fusion of color, texture and depth for robust road scene understanding
abstract
We propose a novel approach to pixel-level semantic labeling, which aims to rapidly infer the coarse layout of street scenes from color, texture and depth information in a joint fashion using a randomized decision forest. The recovered pixellevel class probability maps provide a general purpose basis to guide more elaborate vision algorithms. To demonstrate the richness of our labeling, we extend the well-known Stixel model to use the semantic labels as input cues. In addition, we employ our generated low-level information as an attention mechanism for a vehicle detector. In both cases, recognition performance and accuracy are significantly improved. In our experimental evaluation on the public KITTI benchmark, we thoroughly study the characteristics of different feature channels as well as their contribution to the overall pixellevel labeling result. Our results underline that the combination of several orthogonal feature channels in a joint model is key to superior performance. This performance improvement comes at little additional cost, given that our approach is able to operate at 100 Hz using a GPU implementation.
Timo Scharwächter, Uwe Franke
Intelligent Vehicles Symposium2
2014 Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice
Peter Pinggera, David Pfeiffer, Uwe Franke, Rudolf Mester
ECCV (2)3
2014 Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding
Timo Scharwächter, Markus Enzweiler, Uwe Franke, Stefan Roth 0001
ECCV (5)3
2014 Spider-based Stixel object segmentation
abstract
Stereo vision has established in the field of driver assistance and vehicular safety systems. Next steps along the road towards accident free driving aim to assist the driver in increasingly complex situations such as inner-city traffic. In order to achieve these goals, it is desirable to incorporate higher-order object knowledge in the stereo vision-based understanding of traffic scenes. In particular, object shape and dimension information can help to achieve correct interpretations. However, typically this kind of higher-order information results in a difficult energy minimization problem since large areas of the input image have to be constrained. In this contribution, an efficient global optimization approach based on dynamic programming is proposed that is able to take into account such higher-order object knowledge. The approach is built upon a simple tree representation of the Dynamic Stixel World, an efficient super-pixel object representation. Experiments show that object segmentation can be improved significantly by means of the higher-order object information.
Friedrich Erbs, Andreas Witte, Timo Scharwächter, Rudolf Mester, Uwe Franke
Intelligent Vehicles Symposium5
2014 Will this car change the lane? - Turn signal recognition in the frequency domain
abstract
Understanding the intention of other road users is a key requirement for autonomous driving. In this regard, one particularly relevant cue is a flashing turn signal, since it gives an important hint regarding the intended driving direction of another vehicle in the next few seconds. As such, turn signals can be considered as one of the first methods invented for car-to-car communication. In contrast to modern radio-based approaches, turn signals are installed in almost every vehicle. However, only image-based methods are able to detect, recognize and understand those signals. In this paper, we present a new method to recognize turn signals of other vehicles in images. Our approach builds upon a robust vehicle detector and involves three major steps applied to each detected vehicle: light spot detection, feature extraction through FFT-based analysis of the temporal signal behavior at each detected light spot, and AdaBoost classification of the extracted feature set. In our experiments, we use solely virtually-generated data for training and evaluate the proposed approach on a large 30 minute real-world image sequence. Our results indicate competitive performance at real-time speeds.
Björn Fröhlich, Markus Enzweiler, Uwe Franke
Intelligent Vehicles Symposium3
2014 Visual guard rail detection for advanced highway assistance systems
abstract
In this paper we present a novel method to detect guard rails in highway scenarios using a stereo camera setup. In contrast to previous methods, we combine geometry information with appearance cues using a state-of-the-art feature encoding method. In our system pipeline, we follow a hough-based approach to localize potential guard rails in the image and require each detected line to fulfill linearity in depth as well as certain height expectations. To leverage the appearance information, we exploit an efficient bag-of-features representation that relies on randomized clustering forests. The effectiveness of our approach is demonstrated on a large novel dataset with pixel-level annotations of guard rails in real-world highway scenarios.
Timo Scharwächter, Manuela Schuler, Uwe Franke
Intelligent Vehicles Symposium3
2013 Towards multi-cue urban curb recognition
abstract
This paper presents a multi-cue approach to curb recognition in urban traffic. We propose a novel texture-based curb classifier using local receptive field (LRF) features in conjunction with a multi-layer neural network. This classification module operates on both intensity images and on three-dimensional height profile data derived from stereo vision. We integrate the proposed multi-cue curb classifier as an additional measurement module into a state-of-the-art Kaiman filter-based urban lane recognition system. Our experiments involve a challenging real-world dataset captured in urban traffic with manually labeled ground-truth. We quantify the benefit of the proposed multi-cue curb classifier in terms of the improvement in curb localization accuracy of the integrated system. Our results indicate a 25% reduction of the average curb localization error at real-time processing speeds.
Markus Enzweiler, Pierre Greiner, Carsten Knöppel, Uwe Franke
Intelligent Vehicles Symposium4
2013 From stixels to objects - A conditional random field based approach
abstract
Detection and tracking of moving traffic participants like vehicles, pedestrians or bicycles from a mobile platform using a stereo camera system plays a key role in traffic scene understanding and for future driver assistance and safety systems. To this end, this work presents a Bayesian segmentation approach based on the Dynamic Stixel World, an efficient super-pixel object representation. The existence and state estimation of an (initially) unknown number of moving objects and the detection of stationary background is formulated as a time-recursive energy minimization problem that can be solved in real-time by means of the alpha-expansion multi-class graph cut optimization scheme. In order to handle noise, this approach integrates 3D and motion features as well as spatio-temporal prior knowledge in a probabilistic conditional random field (CRF) framework. An optional fusion step with an additional radar sensor combines the advantages of both measuring instruments and yields superior overall results. The performance and robustness of the presented approach is evaluated quantitatively in various challenging traffic scenes.
Friedrich Erbs, Beate Schwarz, Uwe Franke
Intelligent Vehicles Symposium3
2013 LaneLoc: Lane marking based localization using highly accurate maps
abstract
Precise and robust localization in real-world traffic scenarios is a new challenge arising in the context of autonomous driving and future driver assistance systems. The required precision is in the range of a few centimeters. In urban areas this precision cannot be achieved by standard global navigation satellite systems (GNSS). Our novel approach achieves this requirement using a stereo camera system and a highly accurate map containing curbs and road markings. The maps are created beforehand using an extended sensor setup. GNSS position is used for initialization only and is not required during the localization process. In the paper we present the localization process and provide an evaluation on a test track under known conditions as well as a long term evaluation on approximately 50 km of rural roads, where a precision in centimeter-range is achieved.
Markus Schreiber, Carsten Knöppel, Uwe Franke
Intelligent Vehicles Symposium3
2012 Stixmentation - Probabilistic Stixel based Traffic Scene Labeling
Friedrich Erbs, Beate Schwarz, Uwe Franke
BMVC3
2012 Efficient Stixel-based object recognition
abstract
This paper presents a novel attention mechanism to improve stereo-vision based object recognition systems in terms of recognition performance and computational efficiency at the same time. We utilize the Stixel World, a compact medium-level 3D representation of the local environment, as an early focus-of-attention stage for subsequent system modules. In particular, the search space of computationally expensive pattern classifiers is significantly narrowed down. We explicitly couple the 3D Stixel representation with prior knowledge about the object class of interest, i.e. 3D geometry and symmetry, to precisely focus processing on well-defined local regions that are consistent with the environment model. Experiments are conducted on large real-world datasets captured from a moving vehicle in urban traffic. In case of vehicle recognition as an experimental testbed, we demonstrate that the proposed Stixel-based attention mechanism significantly reduces false positive rates at constant sensitivity levels by up to a factor of 8 over state-of-the-art. At the same time, computational costs are reduced by more than an order of magnitude.
Markus Enzweiler, Matthias Hummel, David Pfeiffer, Uwe Franke
Intelligent Vehicles Symposium4
2012 May I enter the roundabout? A time-to-contact computation based on stereo-vision
abstract
This paper presents a stereo-vision based system for the recognition of dangerous situations at roundabouts. At first, we investigate the necessary field of view and viewing direction using videos taken by a panoramic camera. Using the insights of these tests we build up a stereo-vision system. This system is based on the well established disparity estimation scheme Semi-Global Matching and the recently introduced medium-level representation called Dynamic Stixel-World. A time-to-contact measure is defined that makes explicit use of the roundabouts structural characteristics. Using this measure enables us to create a system for driver warning or possible automated intervention. Our empirical studies reveal that the warning decision correctly mimics human driver decisions.
Maximilian Muffert, Timo Milbich, David Pfeiffer, Uwe Franke
Intelligent Vehicles Symposium4
2012 Improving sub-pixel accuracy for long range stereo
Stefan K. Gehrig, Hernán Badino, Uwe Franke
Comput. Vis. Image Underst.3
2011 Towards a Global Optimal Multi-Layer Stixel Representation of Dense 3D Data
abstract
Dense 3D data as delivered by stereo vision systems, modern laser scanners or timeof-flight cameras such as PMD is a key element for 3D scene understanding. Real-time high-level vision systems require a compact and explicit representation of that data which allows for efficient attention control, object detection, and reasoning. Because man-made environments are dominated by planar horizontal and vertical surfaces we approximate the three dimensional scenery by using sets of thin planar rectangles called Stixels. This medium level representation serves as input for further processing steps and applications. Using this novel representation those are not required to process the large amounts of raw 3D data individually. This reconstruction is addressed by means of a unified probabilistic approach. Dynamic programming allows to incorporate real-world constraints such as perspective ordering and delivers an optimal segmentation with respect to freespace and obstacle information. We present results for both stereo vision data and laser data. The real-time capable approach can also be used to fuse the information of multiple data sources. 1
David Pfeiffer, Uwe Franke
BMVC2
2011 Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes
abstract
We propose and evaluate improvements in motion field estimation in order to cope with challenges in real world scenarios. To build a real-time stereo-based three-dimensional vision system which is able to handle illumination changes, textureless regions and fast moving objects observed by a moving platform, we introduce a new approach to support the variational optical flow computation scheme with stereo and feature information. The improved flow result is then used as input for a temporal integrated robust three-dimensional motion field estimation technique. We evaluate the results of our optical flow algorithm and the resulting three-dimensional motion field against approaches known from literature. Tests on both synthetic realistic and real stereo sequences show that our approach is superior to approaches known from literature with respect to density, accuracy and robustness.
Thomas Müller 0008, Jens Rannacher, Clemens Rabe, Uwe Franke
CVPR4
2011 Moving vehicle detection by optimal segmentation of the Dynamic Stixel World
abstract
The reliable detection of moving objects from a moving observer is one of the most challenging and important tasks for driver assistance and safety systems. Modern sensors such as Lidar, Imaging Radar or Stereo Vision deliver range data plus longitudinal motion (Radar) or even full 3D-motion (space-time vision). Based on this data, moving objects have to be separated from the static background to be able to determine their pose and motion state. Usually, heuristics are applied to cluster the data. In order to find the most probable segmentation, we formulate the task as a hypotheses testing problem that allows taking into account various constraints and assumptions simultaneously. We show that the optimal segmentation can be efficiently found by means of dynamic programming, for an arbitrary number of objects in the scene. In this paper we concentrate on the segmentation of space-time data obtained from stereo image sequences. The vision-based depth and motion information is transferred into so called Stixels, a very compact representation of 3D scenes that can also be applied to Lidar or Radar data. It turns out that our optimal segmentation is more robust w.r.t. noisy and erroneous data.
Friedrich Erbs, Alexander Barth, Uwe Franke
Intelligent Vehicles Symposium3
2011 Towards a closer fusion of active and passive safety: Optical flow-based detection of vehicle side collisions
abstract
In recent years, innovative passive safety concepts have been developed that have the potential to further decrease the number of victims of traffic accidents. Such complex passive safety systems (i.e. the Daimler PRE-SAFE Pulse or inflating metal structures in the vehicle doors) typically require lead times for activation/preparation that are longer than classical crash-detecting acceleration or pressure sensors can offer. These concepts require a close integration with active safety systems in order to allow an early assessment of the type, direction, and severity of an imminent collision. This contribution proposes a prototypical side collision detection system that is based on a monocular camera positioned in the side mirrors. For collision detection, the system fuses the detection results from optical flow and a warped bird's eye view in order to allow a robust system reaction in case of an imminent collision with a dynamic object.
Fridtjof Stein, Uwe Franke
Intelligent Vehicles Symposium3
2011 A temporal filter approach for detection and reconstruction of curbs and road surfaces based on Conditional Random Fields
abstract
A temporal filter approach for real-time detection and reconstruction of curbs and road surfaces from 3D point clouds is presented. Instead of local thresholding, as used in many other approaches, a 3D curb model is extracted from the point cloud. The 3D points are classified to different parts of the model (i.e. road and sidewalk) using a temporally integrated Conditional Random Field (CRF). The parameters of curb and road surface are then estimated from the respectively assigned points, providing a temporal connection via a Kalman filter.
Jan Siegemund, Uwe Franke, Wolfgang Förstner
Intelligent Vehicles Symposium2
2011 Stereoscopic Scene Flow Computation for 3D Motion Understanding
Andreas Wedel, Thomas Brox, Tobi Vaudrey, Clemens Rabe, Uwe Franke, Daniel Cremers
Int. J. Comput. Vis.5
2010 Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time
Clemens Rabe, Thomas Müller 0008, Andreas Wedel, Uwe Franke
ECCV (4)4
2010 Efficient representation of traffic scenes by means of dynamic stixels
abstract
Correlation based stereo vision has proven its power in commercially available driver assistance systems. Recently, real-time dense stereo vision has become available on inexpensive FPGA hardware. In order to manage the huge amount of data, a medium-level representation named “Stixel World” has been proposed for further analysis. In this representation the free space in front of the vehicle is limited by adjacent rectangular sticks of a certain width. Distance and height of each so called stixel are determined by those parts of the obstacle it represents. This Stixel World is a compact but flexible representation of the three-dimensional traffic situation. The underlying model assumption is that objects stand on the ground and have approximately vertical pose with a flat surface. So far, this representation is static since it is computed for each frame independently. Driver assistance, however, is most interested in pose and motion of moving obstacles. For this reason, we introduce tracking of stixels in this paper. Using the 6D-Vision Kalman filter framework, lateral as well as longitudinal motion is estimated for each stixel. That way, the grouping of stixels based on similar motion as well as the detection of moving obstacles turns out to be significantly simplified. The new dynamic Stixel World has proven to be well suited as a common basis for the scene understanding tasks of driver assistance and autonomous systems.
David Pfeiffer, Uwe Franke
Intelligent Vehicles Symposium2
2010 Curb reconstruction using Conditional Random Fields
abstract
This paper presents a generic framework for curb detection and reconstruction in the context of driver assistance systems. Based on a 3D point cloud, we estimate the parameters of a 3D curb model, incorporating also the curb adjacent surfaces, e.g. street and sidewalk. We apply an iterative two step approach. First, the measured 3D points, e.g., obtained from dense stereo vision, are assigned to the curb adjacent surfaces using loopy belief propagation on a Conditional Random Field. Based on this result, we reconstruct the surfaces and in particular the curb. Our system is not limited to straight-line curbs, i.e. it is able to deal with curbs of different curvature and varying height. The proposed algorithm runs in real-time on our demonstrator vehicle and is evaluated in urban real-world scenarios. It yields highly accurate results even for low curbs up to 20m distance.
Jan Siegemund, David Pfeiffer, Uwe Franke, Wolfgang Förstner
Intelligent Vehicles Symposium3
2009 Performance Evaluation of Stereo Algorithms for Automotive Applications
Pascal Steingrube, Stefan K. Gehrig, Uwe Franke
ICVS3
2009 Estimating the Driving State of Oncoming Vehicles From a Moving Platform Using Stereo Vision
abstract
A new image-based approach for fast and robust vehicle tracking from a moving platform is presented. Position, orientation, and full motion state, including velocity, acceleration, and yaw rate of a detected vehicle, are estimated from a tracked rigid 3-D point cloud. This point cloud represents a 3-D object model and is computed by analyzing image sequences in both space and time, i.e., by fusion of stereo vision and tracked image features. Starting from an automated initial vehicle hypothesis, tracking is performed by means of an extended Kalman filter. The filter combines the knowledge about the movement of the rigid point cloud's points in the world with the dynamic model of a vehicle. Radar information is used to improve the image-based object detection at far distances. The proposed system is applied to predict the driving path of other traffic participants and currently runs at 25 Hz (640 times 480 images) on our demonstrator vehicle.
Alexander Barth, Uwe Franke
IEEE Trans. Intell. Transp. Syst.2
2009 B-Spline Modeling of Road Surfaces With an Application to Free-Space Estimation
abstract
We propose a general technique for modeling the visible road surface in front of a vehicle. The common assumption of a planar road surface is often violated in reality. A workaround proposed in the literature is the use of a piecewise linear or quadratic function to approximate the road surface. Our approach is based on representing the road surface as a general parametric B-spline curve. The surface parameters are tracked over time using a Kalman filter. The surface parameters are estimated from stereo measurements in the free space. To this end, we adopt a recently proposed road-obstacle segmentation algorithm to include disparity measurements and the B-spline road-surface representation. Experimental results in planar and undulating terrain verify the increase in free-space availability and accuracy using a flexible B-spline for road-surface modeling.
Andreas Wedel, Hernán Badino, Clemens Rabe, Heidi Loose, Uwe Franke, Daniel Cremers
IEEE Trans. Intell. Transp. Syst.5
2008 Efficient Dense Scene Flow from Sparse or Dense Stereo Data
Andreas Wedel, Clemens Rabe, Tobi Vaudrey, Thomas Brox, Uwe Franke, Daniel Cremers
ECCV (1)5
2007 Improving Stereo Sub-Pixel Accuracy for Long Range Stereo
abstract
Dense stereo algorithms are able to estimate disparities at all pixels including untextured regions. Typically these disparities are evaluated at integer disparity steps. A subsequent sub-pixel interpolation often fails to propagate smoothness constraints on a sub-pixel level. The determination of sub-pixel accurate disparities is an active field of research, however, most sub-pixel estimation algorithms focus on textured image areas in order to show their precision. We propose to increase the sub-pixel accuracy in low- textured regions in three possible ways: First, we present an analysis that shows the benefit of evaluating the disparity space at fractional disparities. Second, we introduce a new disparity smoothing algorithm that preserves depth discontinuities and enforces smoothness on a sub-pixel level. Third, we present a novel stereo constraint (gravitational constraint) that assumes sorted disparity values in vertical direction and guides global algorithms to reduce false matches, especially in low-textured regions. Our goal in this work is to obtain an accurate 3D reconstruction. Large- scale 3D reconstruction will benefit heavily from these sub- pixel refinements, especially with a multi-baseline extension. Results based on semi-global matching , obtained with the above mentioned algorithmic extensions are shown for the Middlebury stereo ground truth data sets. The presented improvements, called ImproveSubPix, turn out to be one of the top-performing algorithms when evaluating the set on a sub-pixel level while being computationally efficient. Additional results are presented for urban scenes. The three improvements are independent of the underlying type of stereo algorithm and can also be applied to sparse stereo algorithms.
Stefan K. Gehrig, Uwe Franke
ICCV2
2002 Fast obstacle detection for urban traffic situations
abstract
The early recognition of potentially harmful traffic situations is an important goal of vision-based driver assistance systems. Pedestrians, in particular children, are highly endangered in inner city traffic. Within the DaimlerChrysler urban traffic assistance (UTA) project, we are using stereo vision and motion analysis in order to manage those situations. The flow/depth constraint combines both methods in an elegant way and leads to a robust and powerful detection scheme. A ball bouncing on the road often implies a child crossing the street. Since balls appear very small in the images of our cameras and can move considerably fast, a special algorithm has been developed to achieve maximum recognition reliability.
Uwe Franke, Stefan Heinrich
IEEE Trans. Intell. Transp. Syst.1
2000 Road recognition in urban environment
Frank Paetzold, Uwe Franke
Image Vis. Comput.2
1992 Spectral Entropy-Activity Classification in Adaptive Transform Coding
abstract
A two-dimensional feature space for block classification which discriminates reliably between blocks that require very different processing is proposed. In combination with a block 'activity' measurement, the introduced 'spectral entropy' feature offers the possibility to stabilize the reconstruction quality of transform coding systems for each processed block on a high level. The classification is valid and useful for threshold-based and zonal coding schemes.>
Rudolf Mester, Uwe Franke
IEEE J. Sel. Areas Commun.2
1989 Top-down image segmentation using object detection and contour relaxation
abstract
A novel segmentation technique that starts with the whole image being a single region is presented. First, an object detection scheme, which marks those locations where local statistics deviate significantly from the overall statistics, provides location and approximate shapes of the major objects (regions) in the scene. Exact boundaries are subsequently obtained by a contour relaxation algorithm, which includes a general model for typical region shapes. Object detection and contour relaxation are repeated recursively until a stable segmentation result is achieved. Segmentation results are presented.>
Til Aach, Uwe Franke, Rudolf Mester
ICASSP2
1987 Selective deconvolution: A new approach to extrapolation and spectral analysis of discrete signals
abstract
This paper describes a new algorithm useful for extrapolation and Fourier analysis of discrete signals that are given by a relative small number of samples. The extrapolation is based on the assumption that the discrete Fourier spectrum shows dominant spectral lines. Involving only FFT, the iterative algorithm is not restricted to one-dimensional signals but can also be applied to higher-dimensional problems. Additional knowledge on the signal like band-limitedness or positivity can easily be taken into account.
Uwe Franke
ICASSP1