EDBT 2026 Demo / reviewers in the wild / expert
Uwe Franke
dblp:56/2085
· DBLP profile ↗
52ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
3D vision · 53% Segmentation and scene understanding · 40% Autonomous driving · 5% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
1.0 | 4 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 Tree-Structured Models for Efficient Multi-Cue Scene Labeling · IEEE Trans. Pattern Anal. Mach. Intell. 2017 The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 |
Computer vision › 3D vision
stereo vision |
0.9 | 4 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice · ECCV (2) 2014 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.2 | 1 | 2016 | The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 |
Computer vision › Segmentation and scene understanding › dense prediction
pixel labeling |
0.2 | 1 | 2016 | The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 |
Computer vision › 3D vision
stereo video dataset |
0.2 | 1 | 2016 | The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 |
Robotics › Autonomous driving
urban scene understanding |
0.2 | 1 | 2016 | The Cityscapes Dataset for Semantic Urban Scene Understanding · CVPR 2016 |
Computer vision › 3D vision
motion estimation |
0.2 | 2 | 2011 | Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011 Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time · ECCV (4) 2010 |
Computer vision › 3D vision
scene flow estimation |
0.2 | 2 | 2011 | Stereoscopic Scene Flow Computation for 3D Motion Understanding · Int. J. Comput. Vis. 2011 Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2014 | Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice · ECCV (2) 2014 |
Computer vision › Segmentation and scene understanding › semantic segmentation › efficient semantic segmentation
real-time semantic segmentation |
0.2 | 1 | 2014 | Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding · ECCV (5) 2014 |
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding |
0.2 | 1 | 2014 | Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding · ECCV (5) 2014 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.2 | 2 | 2008 | Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008 Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007 |
Computer vision › 3D vision › motion estimation
optical flow |
0.1 | 1 | 2011 | Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011 |
Computer vision › 3D vision › motion estimation › optical flow
variational optical flow |
0.1 | 1 | 2011 | Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.1 | 1 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 |
Computer vision › 3D vision › stereo vision
stereo image sequences |
0.1 | 1 | 2010 | Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time · ECCV (4) 2010 |
Computer vision › 3D vision › scene flow estimation
dense scene flow |
0.1 | 1 | 2008 | Efficient Dense Scene Flow from Sparse or Dense Stereo Data · ECCV (1) 2008 |
Computer vision › 3D vision
3d reconstruction |
0.1 | 1 | 2007 | Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007 |
Computer vision › 3D vision › stereo vision › stereo matching
dense stereo matching |
0.1 | 1 | 2007 | Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007 |
Computer vision › 3D vision › stereo vision › stereo matching › disparity refinement
sub-pixel disparity estimation |
0.1 | 1 | 2007 | Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007 |
Computer vision › 3D vision › motion estimation
3d motion estimation |
0.0 | 1 | 2011 | Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenes · CVPR 2011 |
Computer vision › 3D vision › stereo vision › stereo matching
semi-global matching |
0.0 | 1 | 2007 | Improving Stereo Sub-Pixel Accuracy for Long Range Stereo · ICCV 2007 |
Image and video coding › transform coding
adaptive transform coding |
0.0 | 1 | 1992 | Spectral Entropy-Activity Classification in Adaptive Transform Coding · IEEE J. Sel. Areas Commun. 1992 |
Image and video coding
transform coding |
0.0 | 1 | 1992 | Spectral Entropy-Activity Classification in Adaptive Transform Coding · IEEE J. Sel. Areas Commun. 1992 |
Methods — techniques the papers use, named apart from their topics
over-segmentation · 0.4global energy minimization · 0.4fully convolutional network · 0.4tree-structured model · 0.3superpixel · 0.3conditional random field · 0.3weak annotation · 0.2deep learning · 0.2real-time inference · 0.2medium-level model · 0.2spectral entropy · 0.0activity measurement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Learning Stixel-based Instance SegmentationabstractStixels have been successfully applied to a wide range of vision tasks in autonomous driving, recently including instance segmentation. However, due to their sparse occurrence in the image, until now Stixels seldomly served as input for Deep Learning algorithms, restricting their utility for such approaches. In this work we present StixelPointNet, a novel method to perform fast instance segmentation directly on Stixels. By regarding the Stixel representation as unstructured data similar to point clouds, architectures like PointNet are able to learn features from Stixels. We use a bounding box detector to propose candidate instances, for which the relevant Stixels are extracted from the input image. On these Stixels, a PointNet models learns binary segmentations, which we then unify throughout the whole image in a final selection step. StixelPointNet achieves state-of-the-art performance on Stixel-level, is considerably faster than pixel-based segmentation methods, and shows that with our approach the Stixel domain can be introduced to many new 3D Deep Learning tasks. Monty Santarossa, Lukas Schneider, Claudius Zelenka, Lars Schmarje, Reinhard Koch, Uwe Franke |
IV | 6 |
| 2020 | Single-Shot 3D Detection of Vehicles from Monocular RGB Images via Geometrically Constrained Keypoints in Real-TimeabstractIn this paper we propose a novel 3D single-shot object detection method for detecting vehicles in monocular RGB images. Our approach lifts 2D detections to 3D space by predicting additional regression and classification parameters and hence keeping the runtime close to pure 2D object detection. The additional parameters are transformed to 3D bounding box keypoints within the network under geometric constraints. Our proposed method features a full 3D description including all three angles of rotation without supervision by any labeled ground truth data for the object's orientation, as it focuses on certain keypoints within the image plane. While our approach can be combined with any modern object detection framework with only little computational overhead, we exemplify the extension of SSD for the prediction of 3D bounding boxes. We test our approach on different datasets for autonomous driving and evaluate it using the challenging KITTI 3D Object Detection as well as the novel nuScenes Object Detection benchmarks. While we achieve competitive results on both benchmarks we outperform current state-of-the-art methods in terms of speed with more than 20 FPS for all tested datasets and image resolutions. Nils Gählert, Jun-Jun Wan, Nicolas Jourdan 0001, Jan Finkbeiner, Uwe Franke, Joachim Denzler |
IV | 5 |
| 2019 | Beyond Bounding Boxes: Using Bounding Shapes for Real-Time 3D Vehicle Detection from Monocular RGB ImagesabstractThe representation of objects as 2D bounding boxes in monocular RGB images limits the faculty of current computer vision systems to 2D object detection. It fails to provide crucial information such as the orientation of other vehicles, which is vital for autonomous driving. At the same time, real-time performance is essential to qualify an approach for deployment in a productive environment. In order to tackle this problem, we present an approach that predicts several key points selected from a virtual 3D bounding box around a vehicle instead of a pure 2D bounding box. These key points can be interpreted as a bounding shape. With this novel representation we can calculate the actual 3D bounding box of the corresponding object. Thanks to the straightforward implementation of bounding shape in any current state-of-the-art 2D object detector both for singleshot frameworks like YOLO or SSD as well as for two-stage detectors like Faster-RCNN with a minimum of computational overhead, it is able to be run in real-time while providing additional useful information for vehicle detection. We exemplify the extension of SSD to Bounding Shape SSD ( BS3D) and evaluate our approach using the challenging KITTI as well as the novel VIPER dataset. Nils Gählert, Jun-Jun Wan, Michael Weber 0009, Johann Marius Zöllner, Uwe Franke, Joachim Denzler |
IV | 5 |
| 2019 | Slanted Stixels: A Way to Represent Steep StreetsabstractAbstract This work presents and evaluates a novel compact scene representation based on Stixels that infers geometric and semantic information. Our approach overcomes the previous rather restrictive geometric assumptions for Stixels by introducing a novel depth model to account for non-flat roads and slanted objects. Both semantic and depth cues are used jointly to infer the scene representation in a sound global energy minimization formulation. Furthermore, a novel approximation scheme is introduced in order to significantly reduce the computational complexity of the Stixel algorithm, and then achieve real-time computation capabilities. The idea is to first perform an over-segmentation of the image, discarding the unlikely Stixel cuts, and apply the algorithm only on the remaining Stixel cuts. This work presents a novel over-segmentation strategy based on a fully convolutional network, which outperforms an approach based on using local extrema of the disparity map. We evaluate the proposed methods in terms of semantic and geometric accuracy as well as run-time on four publicly available benchmark datasets. Our approach maintains accuracy on flat road scene datasets while improving substantially on a novel non-flat road dataset. Daniel Hernández Juárez, Lukas Schneider, Pau Cebrian, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys, Juan C. Moure |
Int. J. Comput. Vis. | 7 |
| 2018 | MB-Net: MergeBoxes for Real-Time 3D Vehicles DetectionabstractHigh performance vehicle detection and pose esti- mation in RGB images is essential for driver assistance systems as well as for autonomous vehicles. Classical 2D box-based detection schemes allow roughly estimating the position of other vehicles, but not their orientation relative to the ego-vehicle. Recent approaches use 3D models to derive the pose of other vehicles from single monocular images but do not reach real- time performance. In this paper we present an approach that achieves competitive performance on the challenging KITTI Object Detection and orientation Estimation benchmark while being the fastest approach with over 40 FPS. The key is a novel representation named MergeBox whose parameters can be estimated extremely efficiently. We extend SSD-a current fast state-of-the-art 2D box object detector- with this representation to our MB-Net. In contrast to all other current state-of-the-art methods we do not require explicit information on the object orientation for training our model. This reduces label costs significantly, a further advantage for practical applications that require labeling of databases that are much bigger than those used for research. Nils Gählert, Marina Mayer, Lukas Schneider, Uwe Franke, Joachim Denzler |
Intelligent Vehicles Symposium | 4 |
| 2018 | Box2Pix: Single-Shot Instance Segmentation by Assigning Pixels to Object BoxesabstractThe task of semantic instance segmentation has gained a large interest within academia as well as industry, especially in the context of autonomous driving. While several published approaches achieve very strong results, only few of them achieve frame rates that are sufficient for the automotive domain. We present an approach that achieves competitive results on the Cityscapes [1] and KITTI [2] datasets, while being twice as fast as any other existing approach. Our method relies on a single fully-convolutional network (FCN [3]) predicting object bounding boxes, as well as pixel-wise semantic object classes and an offset vector pointing to corresponding object centers. Using those outputs, we present an efficient and simple post-processing that assigns each object pixel to its best matching object detection, resulting in an instance segmentation obtained at real-time speeds. Jonas Uhrig, Eike Rehder, Björn Fröhlich, Uwe Franke, Thomas Brox |
Intelligent Vehicles Symposium | 4 |
| 2017 | Sparsity Invariant CNNsabstractIn this paper, we consider convolutional neural networks operating on sparse inputs with an application to depth completion from sparse laser scan data. First, we show that traditional convolutional networks perform poorly when applied to sparse data even when the location of missing data is provided to the network. To overcome this problem, we propose a simple yet effective sparse convolution layer which explicitly considers the location of missing data during the convolution operation. We demonstrate the benefits of the proposed network architecture in synthetic and real experiments with respect to various baseline approaches. Compared to dense baselines, the proposed sparse convolution network generalizes well to novel datasets and is invariant to the level of sparsity in the data. For our evaluation, we derive a novel dataset from the KITTI benchmark, comprising over 94k depth annotated RGB images. Our dataset allows for training and evaluating depth completion and depth prediction techniques in challenging real-world settings and is available online at: www.cvlibs.net/datasets/kitti. Jonas Uhrig, Nick Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, Andreas Geiger 0001 |
3DV | 4 |
| 2017 | Slanted Stixels: Representing San Francisco's Steepest Streets
Daniel Hernández Juárez, Lukas Schneider, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys |
BMVC | 7 |
| 2017 | Detecting unexpected obstacles for self-driving cars: Fusing deep learning and geometric modelingabstractThe detection of small road hazards, such as lost cargo, is a vital capability for self-driving cars. We tackle this challenging and rarely addressed problem with a vision system that leverages appearance, contextual as well as geometric cues. To utilize the appearance and contextual cues, we propose a new deep learning-based obstacle detection framework. Here a variant of a fully convolutional network is proposed to predict a pixel-wise semantic labeling of (i) free-space, (ii) on-road unexpected obstacles, and (iii) background. The geometric cues are exploited using a state-of-the-art detection approach that predicts obstacles from stereo input images via model-based statistical hypothesis tests. We present a principled Bayesian framework to fuse the semantic and stereo-based detection results. The mid-level Stixel representation is used to describe obstacles in a flexible, compact and robust manner. We evaluate our new obstacle detection system on the Lost and Found dataset, which includes very challenging scenes with obstacles of only 5 cm height. Overall, we report a major improvement over the state-of-the-art, with a performance gain of 27.4%. In particular, we achieve a detection rate of over 90% for distances of up to 50 m. Our system operates at 22 Hz on our self-driving platform. Sebastian Ramos, Stefan K. Gehrig, Peter Pinggera, Uwe Franke, Carsten Rother |
Intelligent Vehicles Symposium | 4 |
| 2017 | RegNet: Multimodal sensor registration using deep neural networksabstractIn this paper, we present RegNet, the first deep convolutional neural network (CNN) to infer a 6 degrees of freedom (DOF) extrinsic calibration between multimodal sensors, exemplified using a scanning LiDAR and a monocular camera. Compared to existing approaches, RegNet casts all three conventional calibration steps (feature extraction, feature matching and global regression) into a single real-time capable CNN. Our method does not require any human interaction and bridges the gap between classical offline and target-less online calibration approaches as it provides both a stable initial estimation as well as a continuous online correction of the extrinsic parameters. During training we randomly decalibrate our system in order to train RegNet to infer the correspondence between projected depth measurements and RGB image and finally regress the extrinsic calibration. Additionally, with an iterative execution of multiple CNNs, that are trained on different magnitudes of decalibration, our approach compares favorably to state-of-the-art methods in terms of a mean calibration error of 0.28° for the rotational and 6 cm for the translation components even for large decalibrations up to 1.5 m and 20°. Nick Schneider, Florian Piewak, Christoph Stiller, Uwe Franke |
Intelligent Vehicles Symposium | 4 |
| 2017 | The Stixel World: A medium-level representation of traffic scenes
Marius Cordts, Timo Rehfeld, Lukas Schneider, David Pfeiffer, Markus Enzweiler, Stefan Roth 0001, Marc Pollefeys, Uwe Franke |
Image Vis. Comput. | 8 |
| 2017 | Stereo vision during adverse weather - Using priors to increase robustness in real-time stereo vision
Stefan K. Gehrig, Nicolai Schneider, Reto Stalder, Uwe Franke |
Image Vis. Comput. | 4 |
| 2017 | Tree-Structured Models for Efficient Multi-Cue Scene LabelingabstractWe propose a novel approach to semantic scene labeling in urban scenarios, which aims to combine excellent recognition performance with highest levels of computational efficiency. To that end, we exploit efficient tree-structured models on two levels: pixels and superpixels. At the pixel level, we propose to unify pixel labeling and the extraction of semantic texton features within a single architecture, so-called encode-and-classify trees. At the superpixel level, we put forward a multi-cue segmentation tree that groups superpixels at multiple granularities. Through learning, the segmentation tree effectively exploits and aggregates a wide range of complementary information present in the data. A tree-structured CRF is then used to jointly infer the labels of all regions across the tree. Finally, we introduce a novel object-centric evaluation method that specifically addresses the urban setting with its strongly varying object scales. Our experiments demonstrate competitive labeling performance compared to the state of the art, while achieving near real-time frame rates of up to 20 fps. Marius Cordts, Timo Rehfeld, Markus Enzweiler, Uwe Franke, Stefan Roth 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | The Cityscapes Dataset for Semantic Urban Scene UnderstandingabstractVisual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately captures the complexity of real-world urban scenes. To address this, we introduce Cityscapes, a benchmark suite and large-scale dataset to train and test approaches for pixel-level and instance-level semantic labeling. Cityscapes is comprised of a large, diverse set of stereo video sequences recorded in streets from 50 different cities. 5000 of these images have high quality pixel-level annotations, 20 000 additional images have coarse annotations to enable methods that leverage large volumes of weakly-labeled data. Crucially, our effort exceeds previous attempts in terms of dataset size, annotation richness, scene variability, and complexity. Our accompanying empirical study provides an in-depth analysis of the dataset characteristics, as well as a performance evaluation of several state-of-the-art approaches based on our benchmark. Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth 0001, Bernt Schiele |
CVPR | 7 |
| 2016 | Visual odometry driven online calibration for monocular LiDAR-camera systemsabstractRecently LiDAR-camera systems have rapidly emerged in many applications. The integration of laser range-finding technologies into existing vision systems enables a more comprehensive understanding of 3D structure of the environment. The advantage, however, relies on a good geometrical calibration between the LiDAR and the image sensors. In this paper we consider visual odometry, a discipline in computer vision and robotics, in the context of recently emerging online sensory calibration studies. By embedding the online calibration problem into a LiDAR-monocular visual odometry technique, the temporal change of extrinsic parameters can be tracked and compensated effectively. Hsiang-Jen Chien, Reinhard Klette, Nick Schneider, Uwe Franke |
ICPR | 4 |
| 2016 | Lost and Found: detecting small road hazards for self-driving vehiclesabstractDetecting small obstacles on the road ahead is a critical part of the driving task which has to be mastered by fully autonomous cars. In this paper, we present a method based on stereo vision to reliably detect such obstacles from a moving vehicle. Peter Pinggera, Sebastian Ramos, Stefan K. Gehrig, Uwe Franke, Carsten Rother, Rudolf Mester |
IROS | 4 |
| 2016 | Semantic Stixels: Depth is not enoughabstractIn this paper we present Semantic Stixels, a novel vision-based scene model geared towards automated driving. Our model jointly infers the geometric and semantic layout of a scene and provides a compact yet rich abstraction of both cues using Stixels as primitive elements. Geometric information is incorporated into our model in terms of pixel-level disparity maps derived from stereo vision. For semantics, we leverage a modern deep learning-based scene labeling approach that provides an object class label for each pixel. Our experiments involve an in-depth analysis and a comprehensive assessment of the constituent parts of our approach using three public benchmark datasets. We evaluate the geometric and semantic accuracy of our model and analyze the underlying run-times and the complexity of the obtained representation. Our results indicate that the joint treatment of both cues on the Semantic Stixel level yields a highly compact environment representation while maintaining an accuracy comparable to the two individual pixel-level input data sources. Moreover, our framework compares favorably to related approaches in terms of computational costs and operates in real-time. Lukas Schneider, Marius Cordts, Timo Rehfeld, David Pfeiffer, Markus Enzweiler, Uwe Franke, Marc Pollefeys, Stefan Roth 0001 |
Intelligent Vehicles Symposium | 6 |
| 2016 | Context-based multi-target tracking with occlusion handling
Junli Tao, Uwe Franke, Reinhard Klette |
Mach. Vis. Appl. | 2 |
| 2015 | What Is in Front? Multiple-Object Detection and Tracking with Dynamic Occlusion Handling
Junli Tao, Markus Enzweiler, Uwe Franke, David Pfeiffer, Reinhard Klette |
CAIP (1) | 3 |
| 2015 | High-performance long range obstacle detection using stereo visionabstractReliable detection of obstacles at long range is crucial for the timely response to hazards by fast-moving safety-critical platforms like autonomous cars. We present a novel method for the joint detection and localization of distant obstacles using a stereo vision system on a moving platform. The approach is applicable to both static and moving obstacles and pushes the limits of detection performance as well as localization accuracy. Peter Pinggera, Uwe Franke, Rudolf Mester |
IROS | 2 |
| 2015 | Low-level fusion of color, texture and depth for robust road scene understandingabstractWe propose a novel approach to pixel-level semantic labeling, which aims to rapidly infer the coarse layout of street scenes from color, texture and depth information in a joint fashion using a randomized decision forest. The recovered pixellevel class probability maps provide a general purpose basis to guide more elaborate vision algorithms. To demonstrate the richness of our labeling, we extend the well-known Stixel model to use the semantic labels as input cues. In addition, we employ our generated low-level information as an attention mechanism for a vehicle detector. In both cases, recognition performance and accuracy are significantly improved. In our experimental evaluation on the public KITTI benchmark, we thoroughly study the characteristics of different feature channels as well as their contribution to the overall pixellevel labeling result. Our results underline that the combination of several orthogonal feature channels in a joint model is key to superior performance. This performance improvement comes at little additional cost, given that our approach is able to operate at 100 Hz using a GPU implementation. Timo Scharwächter, Uwe Franke |
Intelligent Vehicles Symposium | 2 |
| 2014 | Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice
Peter Pinggera, David Pfeiffer, Uwe Franke, Rudolf Mester |
ECCV (2) | 3 |
| 2014 | Stixmantics: A Medium-Level Model for Real-Time Semantic Scene Understanding
Timo Scharwächter, Markus Enzweiler, Uwe Franke, Stefan Roth 0001 |
ECCV (5) | 3 |
| 2014 | Spider-based Stixel object segmentationabstractStereo vision has established in the field of driver assistance and vehicular safety systems. Next steps along the road towards accident free driving aim to assist the driver in increasingly complex situations such as inner-city traffic. In order to achieve these goals, it is desirable to incorporate higher-order object knowledge in the stereo vision-based understanding of traffic scenes. In particular, object shape and dimension information can help to achieve correct interpretations. However, typically this kind of higher-order information results in a difficult energy minimization problem since large areas of the input image have to be constrained. In this contribution, an efficient global optimization approach based on dynamic programming is proposed that is able to take into account such higher-order object knowledge. The approach is built upon a simple tree representation of the Dynamic Stixel World, an efficient super-pixel object representation. Experiments show that object segmentation can be improved significantly by means of the higher-order object information. Friedrich Erbs, Andreas Witte, Timo Scharwächter, Rudolf Mester, Uwe Franke |
Intelligent Vehicles Symposium | 5 |
| 2014 | Will this car change the lane? - Turn signal recognition in the frequency domainabstractUnderstanding the intention of other road users is a key requirement for autonomous driving. In this regard, one particularly relevant cue is a flashing turn signal, since it gives an important hint regarding the intended driving direction of another vehicle in the next few seconds. As such, turn signals can be considered as one of the first methods invented for car-to-car communication. In contrast to modern radio-based approaches, turn signals are installed in almost every vehicle. However, only image-based methods are able to detect, recognize and understand those signals. In this paper, we present a new method to recognize turn signals of other vehicles in images. Our approach builds upon a robust vehicle detector and involves three major steps applied to each detected vehicle: light spot detection, feature extraction through FFT-based analysis of the temporal signal behavior at each detected light spot, and AdaBoost classification of the extracted feature set. In our experiments, we use solely virtually-generated data for training and evaluate the proposed approach on a large 30 minute real-world image sequence. Our results indicate competitive performance at real-time speeds. Björn Fröhlich, Markus Enzweiler, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2014 | Visual guard rail detection for advanced highway assistance systemsabstractIn this paper we present a novel method to detect guard rails in highway scenarios using a stereo camera setup. In contrast to previous methods, we combine geometry information with appearance cues using a state-of-the-art feature encoding method. In our system pipeline, we follow a hough-based approach to localize potential guard rails in the image and require each detected line to fulfill linearity in depth as well as certain height expectations. To leverage the appearance information, we exploit an efficient bag-of-features representation that relies on randomized clustering forests. The effectiveness of our approach is demonstrated on a large novel dataset with pixel-level annotations of guard rails in real-world highway scenarios. Timo Scharwächter, Manuela Schuler, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2013 | Towards multi-cue urban curb recognitionabstractThis paper presents a multi-cue approach to curb recognition in urban traffic. We propose a novel texture-based curb classifier using local receptive field (LRF) features in conjunction with a multi-layer neural network. This classification module operates on both intensity images and on three-dimensional height profile data derived from stereo vision. We integrate the proposed multi-cue curb classifier as an additional measurement module into a state-of-the-art Kaiman filter-based urban lane recognition system. Our experiments involve a challenging real-world dataset captured in urban traffic with manually labeled ground-truth. We quantify the benefit of the proposed multi-cue curb classifier in terms of the improvement in curb localization accuracy of the integrated system. Our results indicate a 25% reduction of the average curb localization error at real-time processing speeds. Markus Enzweiler, Pierre Greiner, Carsten Knöppel, Uwe Franke |
Intelligent Vehicles Symposium | 4 |
| 2013 | From stixels to objects - A conditional random field based approachabstractDetection and tracking of moving traffic participants like vehicles, pedestrians or bicycles from a mobile platform using a stereo camera system plays a key role in traffic scene understanding and for future driver assistance and safety systems. To this end, this work presents a Bayesian segmentation approach based on the Dynamic Stixel World, an efficient super-pixel object representation. The existence and state estimation of an (initially) unknown number of moving objects and the detection of stationary background is formulated as a time-recursive energy minimization problem that can be solved in real-time by means of the alpha-expansion multi-class graph cut optimization scheme. In order to handle noise, this approach integrates 3D and motion features as well as spatio-temporal prior knowledge in a probabilistic conditional random field (CRF) framework. An optional fusion step with an additional radar sensor combines the advantages of both measuring instruments and yields superior overall results. The performance and robustness of the presented approach is evaluated quantitatively in various challenging traffic scenes. Friedrich Erbs, Beate Schwarz, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2013 | LaneLoc: Lane marking based localization using highly accurate mapsabstractPrecise and robust localization in real-world traffic scenarios is a new challenge arising in the context of autonomous driving and future driver assistance systems. The required precision is in the range of a few centimeters. In urban areas this precision cannot be achieved by standard global navigation satellite systems (GNSS). Our novel approach achieves this requirement using a stereo camera system and a highly accurate map containing curbs and road markings. The maps are created beforehand using an extended sensor setup. GNSS position is used for initialization only and is not required during the localization process. In the paper we present the localization process and provide an evaluation on a test track under known conditions as well as a long term evaluation on approximately 50 km of rural roads, where a precision in centimeter-range is achieved. Markus Schreiber, Carsten Knöppel, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2012 | Stixmentation - Probabilistic Stixel based Traffic Scene Labeling
Friedrich Erbs, Beate Schwarz, Uwe Franke |
BMVC | 3 |
| 2012 | Efficient Stixel-based object recognitionabstractThis paper presents a novel attention mechanism to improve stereo-vision based object recognition systems in terms of recognition performance and computational efficiency at the same time. We utilize the Stixel World, a compact medium-level 3D representation of the local environment, as an early focus-of-attention stage for subsequent system modules. In particular, the search space of computationally expensive pattern classifiers is significantly narrowed down. We explicitly couple the 3D Stixel representation with prior knowledge about the object class of interest, i.e. 3D geometry and symmetry, to precisely focus processing on well-defined local regions that are consistent with the environment model. Experiments are conducted on large real-world datasets captured from a moving vehicle in urban traffic. In case of vehicle recognition as an experimental testbed, we demonstrate that the proposed Stixel-based attention mechanism significantly reduces false positive rates at constant sensitivity levels by up to a factor of 8 over state-of-the-art. At the same time, computational costs are reduced by more than an order of magnitude. Markus Enzweiler, Matthias Hummel, David Pfeiffer, Uwe Franke |
Intelligent Vehicles Symposium | 4 |
| 2012 | May I enter the roundabout? A time-to-contact computation based on stereo-visionabstractThis paper presents a stereo-vision based system for the recognition of dangerous situations at roundabouts. At first, we investigate the necessary field of view and viewing direction using videos taken by a panoramic camera. Using the insights of these tests we build up a stereo-vision system. This system is based on the well established disparity estimation scheme Semi-Global Matching and the recently introduced medium-level representation called Dynamic Stixel-World. A time-to-contact measure is defined that makes explicit use of the roundabouts structural characteristics. Using this measure enables us to create a system for driver warning or possible automated intervention. Our empirical studies reveal that the warning decision correctly mimics human driver decisions. Maximilian Muffert, Timo Milbich, David Pfeiffer, Uwe Franke |
Intelligent Vehicles Symposium | 4 |
| 2012 | Improving sub-pixel accuracy for long range stereo
Stefan K. Gehrig, Hernán Badino, Uwe Franke |
Comput. Vis. Image Underst. | 3 |
| 2011 | Towards a Global Optimal Multi-Layer Stixel Representation of Dense 3D DataabstractDense 3D data as delivered by stereo vision systems, modern laser scanners or timeof-flight cameras such as PMD is a key element for 3D scene understanding. Real-time high-level vision systems require a compact and explicit representation of that data which allows for efficient attention control, object detection, and reasoning. Because man-made environments are dominated by planar horizontal and vertical surfaces we approximate the three dimensional scenery by using sets of thin planar rectangles called Stixels. This medium level representation serves as input for further processing steps and applications. Using this novel representation those are not required to process the large amounts of raw 3D data individually. This reconstruction is addressed by means of a unified probabilistic approach. Dynamic programming allows to incorporate real-world constraints such as perspective ordering and delivers an optimal segmentation with respect to freespace and obstacle information. We present results for both stereo vision data and laser data. The real-time capable approach can also be used to fuse the information of multiple data sources. 1 David Pfeiffer, Uwe Franke |
BMVC | 2 |
| 2011 | Feature- and depth-supported modified total variation optical flow for 3D motion field estimation in real scenesabstractWe propose and evaluate improvements in motion field estimation in order to cope with challenges in real world scenarios. To build a real-time stereo-based three-dimensional vision system which is able to handle illumination changes, textureless regions and fast moving objects observed by a moving platform, we introduce a new approach to support the variational optical flow computation scheme with stereo and feature information. The improved flow result is then used as input for a temporal integrated robust three-dimensional motion field estimation technique. We evaluate the results of our optical flow algorithm and the resulting three-dimensional motion field against approaches known from literature. Tests on both synthetic realistic and real stereo sequences show that our approach is superior to approaches known from literature with respect to density, accuracy and robustness. Thomas Müller 0008, Jens Rannacher, Clemens Rabe, Uwe Franke |
CVPR | 4 |
| 2011 | Moving vehicle detection by optimal segmentation of the Dynamic Stixel WorldabstractThe reliable detection of moving objects from a moving observer is one of the most challenging and important tasks for driver assistance and safety systems. Modern sensors such as Lidar, Imaging Radar or Stereo Vision deliver range data plus longitudinal motion (Radar) or even full 3D-motion (space-time vision). Based on this data, moving objects have to be separated from the static background to be able to determine their pose and motion state. Usually, heuristics are applied to cluster the data. In order to find the most probable segmentation, we formulate the task as a hypotheses testing problem that allows taking into account various constraints and assumptions simultaneously. We show that the optimal segmentation can be efficiently found by means of dynamic programming, for an arbitrary number of objects in the scene. In this paper we concentrate on the segmentation of space-time data obtained from stereo image sequences. The vision-based depth and motion information is transferred into so called Stixels, a very compact representation of 3D scenes that can also be applied to Lidar or Radar data. It turns out that our optimal segmentation is more robust w.r.t. noisy and erroneous data. Friedrich Erbs, Alexander Barth, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2011 | Towards a closer fusion of active and passive safety: Optical flow-based detection of vehicle side collisionsabstractIn recent years, innovative passive safety concepts have been developed that have the potential to further decrease the number of victims of traffic accidents. Such complex passive safety systems (i.e. the Daimler PRE-SAFE Pulse or inflating metal structures in the vehicle doors) typically require lead times for activation/preparation that are longer than classical crash-detecting acceleration or pressure sensors can offer. These concepts require a close integration with active safety systems in order to allow an early assessment of the type, direction, and severity of an imminent collision. This contribution proposes a prototypical side collision detection system that is based on a monocular camera positioned in the side mirrors. For collision detection, the system fuses the detection results from optical flow and a warped bird's eye view in order to allow a robust system reaction in case of an imminent collision with a dynamic object. Fridtjof Stein, Uwe Franke |
Intelligent Vehicles Symposium | 3 |
| 2011 | A temporal filter approach for detection and reconstruction of curbs and road surfaces based on Conditional Random FieldsabstractA temporal filter approach for real-time detection and reconstruction of curbs and road surfaces from 3D point clouds is presented. Instead of local thresholding, as used in many other approaches, a 3D curb model is extracted from the point cloud. The 3D points are classified to different parts of the model (i.e. road and sidewalk) using a temporally integrated Conditional Random Field (CRF). The parameters of curb and road surface are then estimated from the respectively assigned points, providing a temporal connection via a Kalman filter. Jan Siegemund, Uwe Franke, Wolfgang Förstner |
Intelligent Vehicles Symposium | 2 |
| 2011 | Stereoscopic Scene Flow Computation for 3D Motion Understanding
Andreas Wedel, Thomas Brox, Tobi Vaudrey, Clemens Rabe, Uwe Franke, Daniel Cremers |
Int. J. Comput. Vis. | 5 |
| 2010 | Dense, Robust, and Accurate Motion Field Estimation from Stereo Image Sequences in Real-Time
Clemens Rabe, Thomas Müller 0008, Andreas Wedel, Uwe Franke |
ECCV (4) | 4 |
| 2010 | Efficient representation of traffic scenes by means of dynamic stixelsabstractCorrelation based stereo vision has proven its power in commercially available driver assistance systems. Recently, real-time dense stereo vision has become available on inexpensive FPGA hardware. In order to manage the huge amount of data, a medium-level representation named “Stixel World” has been proposed for further analysis. In this representation the free space in front of the vehicle is limited by adjacent rectangular sticks of a certain width. Distance and height of each so called stixel are determined by those parts of the obstacle it represents. This Stixel World is a compact but flexible representation of the three-dimensional traffic situation. The underlying model assumption is that objects stand on the ground and have approximately vertical pose with a flat surface. So far, this representation is static since it is computed for each frame independently. Driver assistance, however, is most interested in pose and motion of moving obstacles. For this reason, we introduce tracking of stixels in this paper. Using the 6D-Vision Kalman filter framework, lateral as well as longitudinal motion is estimated for each stixel. That way, the grouping of stixels based on similar motion as well as the detection of moving obstacles turns out to be significantly simplified. The new dynamic Stixel World has proven to be well suited as a common basis for the scene understanding tasks of driver assistance and autonomous systems. David Pfeiffer, Uwe Franke |
Intelligent Vehicles Symposium | 2 |
| 2010 | Curb reconstruction using Conditional Random FieldsabstractThis paper presents a generic framework for curb detection and reconstruction in the context of driver assistance systems. Based on a 3D point cloud, we estimate the parameters of a 3D curb model, incorporating also the curb adjacent surfaces, e.g. street and sidewalk. We apply an iterative two step approach. First, the measured 3D points, e.g., obtained from dense stereo vision, are assigned to the curb adjacent surfaces using loopy belief propagation on a Conditional Random Field. Based on this result, we reconstruct the surfaces and in particular the curb. Our system is not limited to straight-line curbs, i.e. it is able to deal with curbs of different curvature and varying height. The proposed algorithm runs in real-time on our demonstrator vehicle and is evaluated in urban real-world scenarios. It yields highly accurate results even for low curbs up to 20m distance. Jan Siegemund, David Pfeiffer, Uwe Franke, Wolfgang Förstner |
Intelligent Vehicles Symposium | 3 |
| 2009 | Performance Evaluation of Stereo Algorithms for Automotive Applications
Pascal Steingrube, Stefan K. Gehrig, Uwe Franke |
ICVS | 3 |
| 2009 | Estimating the Driving State of Oncoming Vehicles From a Moving Platform Using Stereo VisionabstractA new image-based approach for fast and robust vehicle tracking from a moving platform is presented. Position, orientation, and full motion state, including velocity, acceleration, and yaw rate of a detected vehicle, are estimated from a tracked rigid 3-D point cloud. This point cloud represents a 3-D object model and is computed by analyzing image sequences in both space and time, i.e., by fusion of stereo vision and tracked image features. Starting from an automated initial vehicle hypothesis, tracking is performed by means of an extended Kalman filter. The filter combines the knowledge about the movement of the rigid point cloud's points in the world with the dynamic model of a vehicle. Radar information is used to improve the image-based object detection at far distances. The proposed system is applied to predict the driving path of other traffic participants and currently runs at 25 Hz (640 times 480 images) on our demonstrator vehicle. Alexander Barth, Uwe Franke |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2009 | B-Spline Modeling of Road Surfaces With an Application to Free-Space EstimationabstractWe propose a general technique for modeling the visible road surface in front of a vehicle. The common assumption of a planar road surface is often violated in reality. A workaround proposed in the literature is the use of a piecewise linear or quadratic function to approximate the road surface. Our approach is based on representing the road surface as a general parametric B-spline curve. The surface parameters are tracked over time using a Kalman filter. The surface parameters are estimated from stereo measurements in the free space. To this end, we adopt a recently proposed road-obstacle segmentation algorithm to include disparity measurements and the B-spline road-surface representation. Experimental results in planar and undulating terrain verify the increase in free-space availability and accuracy using a flexible B-spline for road-surface modeling. Andreas Wedel, Hernán Badino, Clemens Rabe, Heidi Loose, Uwe Franke, Daniel Cremers |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2008 | Efficient Dense Scene Flow from Sparse or Dense Stereo Data
Andreas Wedel, Clemens Rabe, Tobi Vaudrey, Thomas Brox, Uwe Franke, Daniel Cremers |
ECCV (1) | 5 |
| 2007 | Improving Stereo Sub-Pixel Accuracy for Long Range StereoabstractDense stereo algorithms are able to estimate disparities at all pixels including untextured regions. Typically these disparities are evaluated at integer disparity steps. A subsequent sub-pixel interpolation often fails to propagate smoothness constraints on a sub-pixel level. The determination of sub-pixel accurate disparities is an active field of research, however, most sub-pixel estimation algorithms focus on textured image areas in order to show their precision. We propose to increase the sub-pixel accuracy in low- textured regions in three possible ways: First, we present an analysis that shows the benefit of evaluating the disparity space at fractional disparities. Second, we introduce a new disparity smoothing algorithm that preserves depth discontinuities and enforces smoothness on a sub-pixel level. Third, we present a novel stereo constraint (gravitational constraint) that assumes sorted disparity values in vertical direction and guides global algorithms to reduce false matches, especially in low-textured regions. Our goal in this work is to obtain an accurate 3D reconstruction. Large- scale 3D reconstruction will benefit heavily from these sub- pixel refinements, especially with a multi-baseline extension. Results based on semi-global matching , obtained with the above mentioned algorithmic extensions are shown for the Middlebury stereo ground truth data sets. The presented improvements, called ImproveSubPix, turn out to be one of the top-performing algorithms when evaluating the set on a sub-pixel level while being computationally efficient. Additional results are presented for urban scenes. The three improvements are independent of the underlying type of stereo algorithm and can also be applied to sparse stereo algorithms. Stefan K. Gehrig, Uwe Franke |
ICCV | 2 |
| 2002 | Fast obstacle detection for urban traffic situationsabstractThe early recognition of potentially harmful traffic situations is an important goal of vision-based driver assistance systems. Pedestrians, in particular children, are highly endangered in inner city traffic. Within the DaimlerChrysler urban traffic assistance (UTA) project, we are using stereo vision and motion analysis in order to manage those situations. The flow/depth constraint combines both methods in an elegant way and leads to a robust and powerful detection scheme. A ball bouncing on the road often implies a child crossing the street. Since balls appear very small in the images of our cameras and can move considerably fast, a special algorithm has been developed to achieve maximum recognition reliability. Uwe Franke, Stefan Heinrich |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2000 | Road recognition in urban environment
Frank Paetzold, Uwe Franke |
Image Vis. Comput. | 2 |
| 1992 | Spectral Entropy-Activity Classification in Adaptive Transform CodingabstractA two-dimensional feature space for block classification which discriminates reliably between blocks that require very different processing is proposed. In combination with a block 'activity' measurement, the introduced 'spectral entropy' feature offers the possibility to stabilize the reconstruction quality of transform coding systems for each processed block on a high level. The classification is valid and useful for threshold-based and zonal coding schemes.> Rudolf Mester, Uwe Franke |
IEEE J. Sel. Areas Commun. | 2 |
| 1989 | Top-down image segmentation using object detection and contour relaxationabstractA novel segmentation technique that starts with the whole image being a single region is presented. First, an object detection scheme, which marks those locations where local statistics deviate significantly from the overall statistics, provides location and approximate shapes of the major objects (regions) in the scene. Exact boundaries are subsequently obtained by a contour relaxation algorithm, which includes a general model for typical region shapes. Object detection and contour relaxation are repeated recursively until a stable segmentation result is achieved. Segmentation results are presented.> Til Aach, Uwe Franke, Rudolf Mester |
ICASSP | 2 |
| 1987 | Selective deconvolution: A new approach to extrapolation and spectral analysis of discrete signalsabstractThis paper describes a new algorithm useful for extrapolation and Fourier analysis of discrete signals that are given by a relative small number of samples. The extrapolation is based on the assumption that the discrete Fourier spectrum shows dominant spectral lines. Involving only FFT, the iterative algorithm is not restricted to one-dimensional signals but can also be applied to higher-dimensional problems. Additional knowledge on the signal like band-limitedness or positivity can easily be taken into account. Uwe Franke |
ICASSP | 1 |