Fabian Flohr

dblp:148/7150 · also Fabian B. Flohr · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-1499-3790ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BikeActions: An Open Platform and Benchmark for Cyclist-Centric VRU Action Recognition
Max Büttner, Kanak Mazumder, Luca Koecher, Mario Finkbeiner, Sebastian Niebler, Fabian Flohr
ICPR (15)6
2026 SatMap: Revisiting Satellite Maps as Prior for Online HD Map Construction
Kanak Mazumder, Fabian Flohr
ICPR (15)2
2026 LIF-Net: LiDAR-Camera Fusion for 3D Human Pose in Urban Scenes
Max Büttner, Erik Schütz, Fabian Flohr
IV3
2026 PlanTRansformer: Unified Prediction and Planning with Goal-conditioned Transformer
Constantin Selzer, Fabian Flohr
IV2
2026 CommandLM: Data driven behavior level descriptor for ego vehicles
Boris Tokic, Constantin Selzer, Fabian Flohr
IV3
2026 HABIT: Human Action Benchmark for Interactive Traffic in CARLA
Mohan Ramesh, Mark Azer, Fabian Flohr
WACV3
2025 BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving
abstract
Autonomous driving has the potential to set the stage for more efficient future mobility, requiring the research domain to establish trust through safe, reliable and transparent driving. Large Language Models (LLMs) possess reasoning capabilities and natural language understanding, presenting the potential to serve as generalized decision-makers for ego-motion planning that can interact with humans and navigate environments designed for human drivers. While this research avenue is promising, current autonomous driving approaches are challenged by combining 3D spatial grounding and the reasoning and language capabilities of LLMs. We introduce BEV-Driver, an LLM-based model for end-to-end closed-loop driving in CARLA that utilizes latent BEV features as perception input. BEVDriver includes a BEV encoder to efficiently process multi-view images and 3D LiDAR point clouds. Within a common latent space, the BEV features are propagated through a Q-Former to align with natural language instructions and passed to the LLM that predicts and plans precise future trajectories while considering navigation instructions and critical scenarios. On the LangAuto benchmark, our model reaches up to 18.9% higher performance on the Driving Score compared to SoTA methods.
Katharina Winter, Mark Azer, Fabian Flohr
IROS3
2025 BEV-LLM: Leveraging Multimodal BEV Maps for Scene Captioning in Autonomous Driving
abstract
Autonomous driving technology has the potential to transform transportation, but its wide adoption depends on the development of interpretable and transparent decision-making systems. Scene captioning, which generates natural language descriptions of the driving environment, plays a crucial role in enhancing transparency, safety, and human-AI interaction. We introduce BEV-LLM, a lightweight model for 3D captioning of autonomous driving scenes. BEV-LLM leverages BEVFusion to combine 3D LiDAR point clouds and multi-view images, incorporating a novel absolute positional encoding for view-specific scene descriptions. Despite using a small 1B parameter base model, BEV-LLM achieves competitive performance on the nuCaption dataset, surpassing state-of-the-art by up to 5% in BLEU scores. Additionally, we release two new datasets — nu-View (focused on environmental conditions and viewpoints) and GroundView (focused on object grounding) — to better assess scene captioning across diverse driving scenarios and address gaps in current benchmarks, along with initial benchmarking results demonstrating their effectiveness.
Felix Brandstätter, Erik Schütz, Katharina Winter, Fabian Flohr
IV4
2025 FAIR-PED: Fairness Evaluation in Pedestrian Detection Using CLIP
abstract
Beyond safety considerations for pedestrians in autonomous driving, fairness in detection promotes societal acceptance and builds public trust. This work introduces FAIR-PED, a novel framework to examine the impact of pedestrian visual attributes on the fair performance of ML-based detection models. The attribute labels are automatically generated using a pedestrian attribute recognition method that leverages contrastive language-image pretraining (CLIP) models, requiring little to no manual annotation. Over 100 pretrained CLIP models are evaluated based on recall balance and accuracy. The ECP dataset is automatically labeled using the best-performing CLIP model for selected attributes. Bootstrap resampling is employed to validate the results. Fairness is measured using the Equal Opportunity Difference (EOD) metric, which compares the recall rates of subgroups. Evaluations reveal detection rate differences across attributes, with insignificant biases toward women and pedestrians carrying bags (EOD = −1.12% and −0.34%, respectively) and a significantly lower detection rate for laterally viewed pedestrians (EOD = +4.65%). Since people crossing the road are often seen as laterally viewed pedestrians, the consistent presence of view bias across all detectors is particularly concerning.
Mohammad Khoshkdahan, Nicholas Kjär, Fabian Flohr
IV3
2024 Walk-the-Talk: LLM driven pedestrian motion generation
abstract
In the field of autonomous driving, a key challenge is the "reality gap": transferring knowledge gained in simulation to real-world settings. Despite various approaches to mitigate this gap, there’s a notable absence of solutions targeting agent behavior generation which are crucial for mimicking spontaneous, erratic, and realistic actions of traffic participants. Recent advancements in Generative AI have enabled the representation of human activities in semantic space and generate real human motion from textual descriptions. Despite current limitations such as modality constraints, motion sequence length, resource demands, and data specificity, there’s an opportunity to innovate and use these techniques in the intelligent vehicles domain. We propose Walk-the-Talk, a motion generator utilizing Large Language Models (LLMs) to produce reliable pedestrian motions for high-fidelity simulators like CARLA. Thus, we contribute to autonomous driving simulations by aiming to scale realistic, diverse long-tail agent motion data - currently a gap in training datasets. We employ Motion Capture (MoCap) techniques to develop the Walk-the-Talk dataset, which illustrates a broad spectrum of pedestrian behaviors in street-crossing scenarios, ranging from standard walking patterns to extreme behaviors such as drunk walking and near-crash incidents. By utilizing this new dataset within a LLM, we facilitate the creation of realistic pedestrian motion sequences, a capability previously unattainable (cf. Figure 1). Additionally, our findings demonstrate that leveraging the Walk-the-Talk dataset enhances cross-domain generalization and significantly improves the Fréchet Inception Distance (FID) score by approximately 15% on the HumanML3D dataset. https://iv.ee.hm.edu/publications/w-the-t/
Mohan Ramesh, Fabian Flohr
IV2
2023 Weakly Supervised Multi-Modal 3D Human Body Pose Estimation for Autonomous Driving
abstract
Accurate 3D human pose estimation (3D HPE) is crucial for enabling autonomous vehicles (AVs) to make informed decisions and respond proactively in critical road scenarios. Promising results of 3D HPE have been gained in several domains such as human-computer interaction, robotics, sports and medical analytics, often based on data collected in well-controlled laboratory environments. Nevertheless, the transfer of 3D HPE methods to AVs has received limited research attention, due to the challenges posed by obtaining accurate 3D pose annotations and the limited suitability of data from other domains.We present a simple yet efficient weakly supervised approach for 3D HPE in the AV context by employing a high-level sensor fusion between camera and LiDAR data. The weakly supervised setting enables training on the target datasets without any 2D / 3D keypoint labels by using an off-the-shelf 2D joint extractor and pseudo labels generated from LiDAR to image projections. Our approach outperforms state-of-the-art results by up to ~ 13% on the Waymo Open Dataset in the weakly supervised setting and achieves state-of-the-art results in the supervised setting.
Peter Bauer, Arij Bouazizi, Ulrich Kressel, Fabian Flohr
IV4
2022 Point Cloud Generation with Continuous Conditioning
abstract
Generative models can be used to synthesize 3D objects of high quality and diversity. However, there is typically no control over the properties of the generated object.This paper proposes a novel generative adversarial network (GAN) setup that generates 3D point cloud shapes conditioned on a continuous parameter. In an exemplary application, we use this to guide the generative process to create a 3D object with a custom-fit shape. We formulate this generation process in a multi-task setting by using the concept of auxiliary classifier GANs. Further, we propose to sample the generator label input for training from a kernel density estimation (KDE) of the dataset. Our ablations show that this leads to significant performance increase in regions with few samples. Extensive quantitative and qualitative experiments show that we gain explicit control over the object dimensions while maintaining good generation quality and diversity.
Larissa T. Triess, Andre Bühler, David Peter, Fabian Flohr, Johann Marius Zöllner
AISTATS4
2021 VRU Pose-SSD: Multiperson Pose Estimation For Automated Driving
abstract
We present a fast and efficient approach for joint person detection and pose estimation optimized for automated driving (AD) in urban scenarios. We use a multitask weight sharing architecture to jointly train detection and pose estimation. This modular architecture allows us to accommodate different downstream tasks in the future. By systematic large-scale experiments on the Tsinghua-Daimler Urban Pose Dataset (TDUP), we obtain multiple models with varying accuracy-speed trade-offs. We then quantize and optimize our network for deployment and present a detailed analysis of the efficacy of the algorithm. We introduce a two-stage evaluation strategy, which is more suitable for AD and achieve a significant performance improvement in comparison to state-of-the-art approaches. Our optimized model runs at 52~fps on full HD images and still reaches a competitive performance of 32.25~LAMR. We are confident that our work serves as an enabler to tackle higher-level tasks like VRU intention estimation and gesture recognition, which rely on stable pose estimates and will play a crucial role in future AD systems.
Jayanth Ramesh, Bodhisattwa Chakraborty, Renjith Raman, Christoph Weinrich, Anurag Mundhada, Arjun Jain, Fabian Flohr
AAAI8
2021 Simple Pair Pose - Pairwise Human Pose Estimation in Dense Urban Traffic Scenes
abstract
Despite the success of deep learning, human pose estimation remains a challenging problem in particular in dense urban traffic scenarios. Its robustness is important for followup tasks like trajectory prediction and gesture recognition. We are interested in human pose estimation in crowded scenes with overlapping pedestrians, in particular pairwise constellations. We propose a new top-down method that relies on pairwise detections as input and jointly estimates the two poses of such pairs in a single forward pass within a deep convolutional neural network. As availability of automotive datasets providing poses and a fair amount of crowded scenes is limited, we extend the EuroCity Persons dataset by additional images and pose annotations. With 46,975 images and poses of 279,329 persons our new EuroCity Persons Dense Pose dataset is the largest pose dataset recorded from a moving vehicle. In our experiments using this dataset we show improved performance for poses of pedestrian pairs in comparison with a state of the art method for human pose estimation in crowds.
Markus Braun 0003, Fabian Flohr, Sebastian Krebs, Ulrich Kreße, Dariu Gavrila
IV2
2021 UrbanPose: A New Benchmark for VRU Pose Estimation in Urban Traffic Scenes
abstract
Human pose, serving as a robust appearance-invariant mid-level feature, has proven to be effective and efficient for human action recognition and intention estimation. Pose features also have a great potential to improve trajectory prediction for the Vulnerable Road User (VRU) in ADAS or automated driving applications. However, the lack of highly diverse and large VRU pose datasets makes a transfer and application to the VRU rather difficult. This paper introduces the Tsinghua-Daimler Urban Pose dataset (TDUP), a large-scale 2D VRU pose image dataset collected in Chinese urban traffic environments from on-board a moving vehicle. The TDUP dataset contains 21k images with more than 90k high-quality, manually labeled VRU bounding boxes with pose keypoint annotations and additional tags. We optimize four state-of-the-art deep learning approaches (AlphaPose, Mask R-CNN, Pose-SSD and PitPaf) to serve as baselines for the new pose estimation benchmark. We further analyze the effect of using large pre-training datasets and different data proportions as well as optional labeled information during training. Our new benchmark is expected to lay the foundation for further VRU pose studies and to empower the development of accurate VRU trajectory prediction methods in complex urban traffic scenes. The dataset (including an evaluation server) is available on www.urbanpose-dataset.com for non-commercial scientific use.
Diange Yang, Baofeng Wang, Zijie Guo, Rishabh Verma, Jayanth Ramesh, Christoph Weinrich, Ulrich Kressel, Fabian Flohr
IV9
2021 Pose-Guided Person Image Synthesis for Data Augmentation in Pedestrian Detection
abstract
In this paper, we present a data augmentation framework for pedestrian detection using a pose-guided person image synthesis model. The proposed framework can boost the performance of state-of-the-art pedestrian detectors by generating new and unseen pedestrian training samples with controllable appearances and poses. This is achieved by a new latent-consistent adversarial variational auto-encoder (LAVAE) model, leveraging the advantages of conditional variational auto-encoders and conditional generative adversarial networks to disengage and reconstruct person images conditioned on target poses. An additional latent regression path is introduced to preserve appearance information and to guarantee a spatial alignment during transfer. LAVAE goes beyond existing works in restoring structural information and perceptual details with limited annotations and can further benefit the pedestrian detection task in automated driving scenarios. Extensive pedestrian detection and person image synthesis experiments are performed on the EuroCity Person dataset. We show that data augmentation using LAVAE improves the accuracy of state-of-the-art pedestrian detectors significantly. Furthermore, a competitive performance can be observed when we compare LAVAE with other generative models for person image synthesis.
Rong Zhi, Zijie Guo, Wuqiang Zhang, Baofeng Wang, Vitali Kaiser, Julian Wiederer, Fabian Flohr
IV7
2020 Traffic Police Gesture Recognition by Pose Graph Convolutional Networks
abstract
Gestures from traffic police give the authorized information, especially in some urgent situation. Thus, understanding of traffic police instruction accurately and promptly is particularly crucial for the automated driving system. However, this task is a great challenge not only because of the dynamic and diversity characteristics of the human gesture, but also the high requirement for real-time performance in each frame. We propose an online activity recognition method based on pose estimation and Graph Convolutional Networks (GCN) to recognize the traffic police gesture in frame level. The main contribution in this work is the development of an online framework based on graph convolutional networks for traffic police recognition. Our approach obtained the state-of-the-art results on Traffic Police Gesture Recognition (TPGR) dataset.
Zhijie Fang, Wuqiang Zhang, Zijie Guo, Rong Zhi, Baofeng Wang, Fabian Flohr
IV6
2020 Generative Model based Data Augmentation for Special Person Classification
abstract
Big data leads to a great success of deep learning in computer vision. Unfortunately, big datasets are often not balanced in all dimensions and rare cases are often underrepresented. On-board data collection by a moving vehicle can capture thousands of normal pedestrians and vehicles, but what about special persons like police officers, road workers, and school guards? Not only that those types of classes are hard to get, they are crucial to be recognized and classified as such for the task of automated driving. Future self-driving cars need to interact with their environment and need to also understand and follow the signals and instructions of those special persons. In this paper, we show how to classify special person types using Convolutional Neural Networks. The big data imbalance is handled by data augmentation using Generative Models, showing a clear advantage over classical data augmentation. The classification performance of special persons can be significantly improved using our Generative Model based Data Augmentation.
Zijie Guo, Rong Zhi, Wuqiang Zhang, Baofeng Wang, Zhijie Fang, Vitali Kaiser, Julian Wiederer, Fabian Flohr
IV8
2019 Recurrent Neural Network Architectures for Vulnerable Road User Trajectory Prediction
abstract
We present an experimental study comparing various Recurrent Neural Network architectures for the task of Vulnerable Road User (VRU) motion trajectory prediction in the intelligent vehicle domain. Making use of temporal motion cues and visual appearance features, we design multi-cue RNN-based architectures with dedicated optimization process to predict future moving trajectories from historical consecutive frames. Experiments are performed on image sequences recorded from on-board a moving vehicle and public tracking datasets. In particular, the Tsinghua-Daimler Cyclist Benchmark (TDCB) has been augmented with additional annotations (vari-ous VRU types) to support the evaluation of object tracking approaches and trajectory prediction methods. This newly introduced dataset is termed TDCB-Track. We demonstrate the effectiveness of the proposed RNN architectures on the public MOT16 dataset and the TDCB-Track dataset. We show that the proposed approaches outperform simpler baseline methods and stay ahead with the state-of-the-art.
Hui Xiong 0006, Fabian Flohr, Baofeng Wang, Jianqiang Wang 0003, Keqiang Li 0002
IV2
2019 Context-Based Path Prediction for Targets with Switching Dynamics
abstract
Anticipating future situations from streaming sensor data is a key perception challenge for mobile robotics and automated vehicles. We address the problem of predicting the path of objects with multiple dynamic modes. The dynamics of such targets can be described by a Switching Linear Dynamical System (SLDS). However, predictions from this probabilistic model cannot anticipate when a change in dynamic mode will occur. We propose to extract various types of cues with computer vision to provide context on the target’s behavior, and incorporate these in a Dynamic Bayesian Network (DBN). The DBN extends the SLDS by conditioning the mode transition probabilities on additional context states. We describe efficient online inference in this DBN for probabilistic path prediction, accounting for uncertainty in both measurements and target behavior. Our approach is illustrated on two scenarios in the Intelligent Vehicles domain concerning pedestrians and cyclists, so-called Vulnerable Road Users (VRUs). Here, context cues include the static environment of the VRU, its dynamic environment, and its observed actions. Experiments using stereo vision data from a moving vehicle demonstrate that the proposed approach results in more accurate path prediction than SLDS at the relevant short time horizon (1 s). It slightly outperforms a computationally more demanding state-of-the-art method.
Julian F. P. Kooij, Fabian Flohr, Ewoud A. I. Pool, Dariu Gavrila
Int. J. Comput. Vis.2
2019 EuroCity Persons: A Novel Benchmark for Person Detection in Traffic Scenes
abstract
Big data has had a great share in the success of deep learning in computer vision. Recent works suggest that there is significant further potential to increase object detection performance by utilizing even bigger datasets. In this paper, we introduce the EuroCity Persons dataset, which provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes. The images for this dataset were collected on-board a moving vehicle in 31 cities of 12 European countries. With over 238200 person instances manually labeled in over 47300 images, EuroCity Persons is nearly one order of magnitude larger than datasets used previously for person detection in traffic scenes. The dataset furthermore contains a large number of person orientation annotations (over 211200). We optimize four state-of-the-art deep learning approaches (Faster R-CNN, R-FCN, SSD and YOLOv3) to serve as baselines for the new object detection benchmark. We analyze the generalization capabilities of these detectors when trained with the new dataset. We furthermore study the effect of the training set size, the dataset diversity (day- vs. night-time, geographical region), the dataset detail (i.e. availability of object orientation information) and the annotation quality on the detector performance. Finally, we analyze error sources and discuss the road ahead.
Markus Braun 0003, Sebastian Krebs, Fabian Flohr, Dariu Gavrila
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 A Unified Framework for Concurrent Pedestrian and Cyclist Detection
abstract
Extensive research interest has been focused on protecting vulnerable road users in recent years, particularly pedestrians and cyclists, due to their attributes of vulnerability. However, comparatively little effort has been spent on detecting pedestrian and cyclist together, particularly when it concerns quantitative performance analysis on large datasets. In this paper, we present a unified framework for concurrent pedestrian and cyclist detection, which includes a novel detection proposal method (termed UB-MPR) to output a set of object candidates, a discriminative deep model based on Fast R-CNN for classification and localization, and a specific postprocessing step to further improve detection performance. Experiments are performed on a new pedestrian and cyclist dataset containing 30 490 annotated pedestrian and 26 771 cyclist instances in over 50 000 images, recorded from a moving vehicle in the urban traffic of Beijing. Experimental results indicate that the proposed method outperforms other state-of-the-art methods significantly.
Lingxi Li 0001, Fabian Flohr, Jianqiang Wang 0003, Hui Xiong 0006, Bernhard Morys, Shuyue Pan, Dariu Gavrila, Keqiang Li 0002
IEEE Trans. Intell. Transp. Syst.3
2016 A new benchmark for vision-based cyclist detection
abstract
Significant progress has been achieved over the past decade on vision-based pedestrian detection; this has led to active pedestrian safety systems being deployed in most mid- to high-range cars on the market. Comparatively little effort has been spent on vision-based cyclist detection, especially when it concerns quantitative performance analysis on large datasets. We present a large-scale experimental study on cyclist detection where we examine the currently most promising object detection methods; we consider Aggregated Channel Features, Deformable Part Models and Region-based Convolutional Neural Networks. We also introduce a new method called Stereo-Proposal based Fast R-CNN (SP-FRCN) to detect cyclists based on stereo proposals and Fast R-CNN (FRCN) framework. Experiments are performed on a dataset containing 22161 annotated cyclist instances in over 30000 images, recorded from a moving vehicle in the urban traffic of Beijing. Results indicate that all the three solution families can reach top performance around 0.89 average precision on the easy case, but the performance drops gradually with the difficulty increasing. The dataset including rich annotations, stereo images and evaluation scripts (termed “Tsinghua-Daimler Cyclist Benchmark”) is made public to the scientific community, to serve as a common point of reference for future research.
Fabian Flohr, Hui Xiong 0006, Markus Braun 0003, Shuyue Pan, Keqiang Li 0002, Dariu Gavrila
Intelligent Vehicles Symposium2
2016 Driver and pedestrian awareness-based collision risk analysis
abstract
We present a novel approach for vehicle-pedestrian collision risk analysis that incorporates mutual situational awareness, a degree of potential motion coupling and the spatial layout of the environment. The approach uses a Dynamic Bayesian Network (DBN) for modeling the individual object paths; collision risk is subsequently computed by an intersection operation. More specifically, the proposed DBN consists of two subgraphs for modeling pedestrian and vehicle path, respectively. They consist of latent states on top of Switching Linear Dynamical Systems (SLDSs) to anticipate changes in object dynamics. The pedestrian and vehicle-related sub-graphs contain latent states to model whether the pedestrian has seen the oncoming vehicle, and conversely, whether the driver has seen the pedestrian (associated measurements involve the respective head orientations). The pedestrian-related sub-graph furthermore contains a latent state modeling whether the pedestrian is at the curbside or not. Finally, a latent state is shared by the two sub-graphs, which models the potential motion coupling (i.e. at full awareness of the other traffic participant). We consider the scenario of a crossing pedestrian, who might stop or continue walking at the curb, in combination with an approaching vehicle, that might stop or continue driving. In experiments we illustrate that with the proposed approach, a more anticipatory driver warning and/or vehicle control strategy can be implemented.
Markus Roth, Fabian Flohr, Dariu Gavrila
Intelligent Vehicles Symposium2
2015 A Probabilistic Framework for Joint Pedestrian Head and Body Orientation Estimation
abstract
We present a probabilistic framework for the joint estimation of pedestrian head and body orientation from a mobile stereo vision platform. For both head and body parts, we convert the responses of a set of orientation-specific detectors into a (continuous) probability density function. The parts are localized by means of apictorial structureapproach, which balances part-based detector responses with spatial constraints. Head and body orientations are estimated jointly to account for anatomical constraints. The joint single-frame orientation estimates are integrated over time by particle filtering. The experiments involved data from a vehicle-mounted stereo vision camera in a realistic traffic setting; 65 pedestrian tracks were supplied by a state-of-the-art pedestrian tracker. We show that the proposed joint probabilistic orientation estimation framework reduces the mean absolute head and body orientation error up to 15° compared with simpler methods. This results in a mean absolute head/body orientation error of about 21°/19°, which remains fairly constant up to a distance of 25 m. Our system currently runs in near real time (8–9 Hz).
Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila
IEEE Trans. Intell. Transp. Syst.1
2014 Context-Based Pedestrian Path Prediction
Julian F. P. Kooij, Nicolas Schneider, Fabian Flohr, Dariu Gavrila
ECCV (6)3
2014 Joint probabilistic pedestrian head and body orientation estimation
abstract
We present an approach for the joint probabilistic estimation of pedestrian head and body orientation in the context of intelligent vehicles. For both, head and body, we convert the output of a set of orientation-specific detectors into a full (continuous) probability density function. The parts are localized with a pictorial structure approach which balances part-based detector output with spatial constraints. Head and body orientation estimates are furthermore coupled probabilistically to account for anatomical constraints. Finally, the coupled single-frame orientation estimates are integrated over time by particle filtering. The experiments involve 37 pedestrian tracks obtained from an external stereo vision-based pedestrian detector in realistic traffic settings. We show that the proposed joint probabilistic orientation estimation approach reduces the mean head and body orientation error by 10 degrees and more.
Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila
Intelligent Vehicles Symposium1
2013 PedCut: an iterative framework for pedestrian segmentation combining shape models and multiple data cues
abstract
This paper presents an iterative, EM-like framework for accurate pedestrian segmentation, combining generative shape models and multiple data cues.In the E-step, shape priors are introduced in the unary terms of a Conditional Random Field (CRF) formulation, joining other data terms derived from color, texture and disparity cues.In the M-step, the resulting segmentation is used to adapt an Active Shape Model (ASM), after which the EM process alternates.Experiments on the public Penn-Fudan pedestrian dataset suggest that our method outperforms the state-of-the-art.We further provide results on a new Daimler pedestrian dataset, captured from on-board a vehicle, which includes disparity data.This dataset is made public to facilitate benchmarking.
Fabian Flohr, Dariu Gavrila
BMVC1