EDBT 2026 Demo / reviewers in the wild / expert
Dariu Gavrila
dblp:g/DariuGavrila · also Dariu M. Gavrila
· DBLP profile ↗
74ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0002-1810-4196ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust LIDAR Sensor-to-Sensor Domain Generalization via Accumulated-Point Scene Completion
Christoph Rist, Larissa T. Triess, Markus Enzweiler, Dariu Gavrila |
IV | 4 |
| 2025 | SAM-Maps: Road Map Generation for Automated Vehicles in Urban AreasabstractAutomated Vehicles (AVs) rely on up-to-date map information to inform trajectory prediction and planning modules, but these maps are expensive to obtain and update as they are usually annotated by humans. We propose SAM-Maps, a method for automatically generating road maps from aerial images of urban areas that takes advantage of the power of foundation models, requiring no human annotation or additional training to map unseen areas. This method extracts a coarse road graph from the images and then estimates the geometry of the roads from this graph. We evaluate our model on the challenging road layouts of the recent View-of-Delft Prediction dataset by comparing the maps generated using our model to the human-annotated maps, achieving an IoU of 33.3% with our automatic method and an IoU of 56.1% with some human corrections in our method. We also evaluate a trajectory prediction model on our maps to test whether they are sufficiently accurate for downstream tasks. The performance of this model using the map from our automatic method is 37.9% better on the minADE6 metric than not using map data as input. To the best of our knowledge, this is the first method that extracts both the drivable area and road connections of European urban areas from aerial images. The code will be publicly released for research purposes. Matthijs P. van Andel, Hidde J.-H. Boekema, Dariu Gavrila |
IV | 3 |
| 2025 | A Vehicle System for Navigating Among Vulnerable Road Users Including Remote OperationabstractWe present a vehicle system capable of navigating safely and efficiently around Vulnerable Road Users (VRUs), such as pedestrians and cyclists. The system comprises key modules for environment perception, localization and mapping, motion planning, and control, integrated into a prototype vehicle. A key innovation is a motion planner based on Topology-driven Model Predictive Control (T-MPC). The guidance layer generates multiple trajectories in parallel, each representing a distinct strategy for obstacle avoidance or non-passing. The underlying trajectory optimization constrains the joint probability of collision with VRUs under generic uncertainties. To address extraordinary situations (“edge cases”) that go beyond the autonomous capabilities — such as construction zones or encounters with emergency responders — the system includes an option for remote human operation, supported by visual and haptic guidance. In simulation, our motion planner outperforms three baseline approaches in terms of safety and efficiency. We also demonstrate the full system in prototype vehicle tests on a closed track, both in autonomous and remotely operated modes. Oscar de Groot, Alberto Bertipaglia, Hidde J.-H. Boekema, Vishrut Jain, Marcell Kegl, Varun Kotian, Ted de Vries Lentsch, Yancon Lin, Chrysovalanto Messiou, Emma Schippers, Farzam Tajdari, Zimin Xia, Mubariz Zaffar, Ronald M. Ensing, Mario Garzon, Javier Alonso-Mora, Holger Caesar, Laura Ferranti, Riender Happee, Julian F. P. Kooij, Barys N. Shyrokau, Dariu Gavrila |
IV | 24 |
| 2025 | Camera-and LiDAR-based Person Re-IdentificationabstractIn this paper, we introduce a novel method for creating appearance embeddings to identify individual persons using an object re-identification (ReID) framework. We present CLFormer (Camera LiDAR Transformer), a transformer-based architecture that incorporates multi-modal data from both camera and LiDAR sensors. We introduce the 3D Cuboid-Inclusive Point Embedding (3D-CIPE), which leverages rich data from LiDAR point clouds and 3D cuboids to add a learnable embedding into the transformer structure. Additionally, through ablation studies, we explore and analyze various strategies for the early and late fusion of multi-modal input data. To evaluate our proposed CLFormer, we reinterpret the nuScenes dataset [1] for ReID purposes and use it for our experiments. Our method demonstrates a significant improvement in performance, outperforming the image-only baseline with an increase of 2.3 in mean Average Precision (mAP). Sebastian Krebs, Dariu Gavrila |
IV | 2 |
| 2025 | Topology-Driven Parallel Trajectory Optimization in Dynamic EnvironmentsabstractGround robots navigating in complex, dynamic environments must compute collision-free trajectories to avoid obstacles safely and efficiently. Nonconvex optimization is a popular method to compute a trajectory in real time. However, these methods often converge to locally optimal solutions and frequently switch between different local minima, leading to inefficient and unsafe robot motion. In this work, we propose a novel topology-driven trajectory optimization strategy for dynamic environments that plans multiple distinct evasive trajectories to enhance the robot's behavior and efficiency. A global planner iteratively generates trajectories in distinct homotopy classes. These trajectories are then optimized by local planners working in parallel. While each planner shares the same navigation objectives, they are locally constrained to a specific homotopy class, meaning each local planner attempts a different evasive maneuver. The robot then executes the feasible trajectory with the lowest cost in a receding horizon manner. We demonstrate on a mobile robot navigating among pedestrians that our approach leads to faster trajectories than existing planners. Oscar de Groot, Laura Ferranti, Dariu Gavrila, Javier Alonso-Mora |
IEEE Trans. Robotics | 3 |
| 2024 | Multimodal Object Query Initialization for 3D Object Detectionabstract3D object detection models that exploit both LiDAR and camera sensor features are top performers in large-scale autonomous driving benchmarks. A transformer is a popular network architecture used for this task, in which so-called object queries act as candidate objects. Initializing these object queries based on current sensor inputs is a common practice. For this, existing methods strongly rely on LiDAR data however, and do not fully exploit image features. Besides, they introduce significant latency. To overcome these limitations we propose EfficientQ3M, an efficient, modular, and multimodal solution for object query initialization for transformer-based 3D object detection models. The proposed initialization method is combined with a "modality-balanced" transformer decoder where the queries can access all sensor modalities throughout the decoder. In experiments, we outperform the state of the art in transformer-based LiDAR object detection on the competitive nuScenes benchmark and showcase the benefits of input-dependent multimodal query initialization, while being more efficient than the available alternatives for LiDAR-camera initialization. The proposed method can be applied with any combination of sensor modalities as input, demonstrating its modularity. Mathijs R. van Geerenstein, Felicia Ruppel, Klaus Dietmayer, Dariu Gavrila |
ICRA | 4 |
| 2024 | UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-ClassesabstractUnsupervised 3D object detection methods have emerged to leverage vast amounts of data without requiring manual labels for training. Recent approaches rely on dynamic objects for learning to detect mobile objects but penalize the detections of static instances during training. Multiple rounds of (self) training are used to add detected static instances to the set of training targets; this procedure to improve performance is computationally expensive. To address this, we propose the method UNION. We use spatial clustering and self-supervised scene flow to obtain a set of static and dynamic object proposals from LiDAR. Subsequently, object proposals' visual appearances are encoded to distinguish static objects in the foreground and background by selecting static instances that are visually similar to dynamic objects. As a result, static and dynamic mobile objects are obtained together, and existing detectors can be trained with a single training. In addition, we extend 3D object discovery to detection by using object appearance-based cluster labels as pseudo-class labels for training object classification. We conduct extensive experiments on the nuScenes dataset and increase the state-of-the-art performance for unsupervised 3D object discovery, i.e. UNION more than doubles the average precision to 38.4. The code is available at github.com/TedLentsch/UNION. Ted de Vries Lentsch, Holger Caesar, Dariu Gavrila |
NeurIPS | 3 |
| 2024 | EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in TrafficabstractWe present the EuroCity Persons (ECP) 2.0 dataset, a novel image dataset for person detection, tracking and prediction in traffic. The dataset was collected on-board a vehicle driving through 29 cities in 11 European countries. It contains more than 250K unique person trajectories, in more than 2.0M images and comes with a size of 11 TB. ECP2.0 is about one order of magnitude larger than previous state-of-the-art person datasets in automotive context. It offers remarkable diversity in terms of geographical coverage, time of day, weather and seasons. We discuss the novel semi-supervised approach that was used to generate the temporally dense pseudo ground-truth (i.e., 2D bounding boxes, 3D person locations) from sparse, manual annotations at keyframes. Our approach leverages auxiliary LiDAR data for 3D uplifting and vehicle inertial sensing for ego-motion compensation. It incorporates keyframe information in a three-stage approach (tracklet generation, tracklet merging into tracks, track smoothing) for obtaining accurate person trajectories. We validate our pseudo ground-truth generation approach in ablation studies, and show that it significantly outperforms existing methods. Furthermore, we demonstrate its benefits for training and testing of state-of-the-art tracking methods. Our approach provides a speed-up factor of about 34 compared to frame-wise manual annotation. The ECP2.0 dataset is made freely available for non-commercial research use. Sebastian Krebs, Markus Braun 0003, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Hidden Gems: 4D Radar Scene Flow Learning Using Cross-Modal SupervisionabstractThis work proposes a novel approach to 4D radar-based scene flow estimation via cross-modal learning. Our approach is motivated by the co-located sensing redundancy in modern autonomous vehicles. Such redundancy implicitly provides various forms of supervision cues to the radar scene flow estimation. Specifically, we introduce a multi-task model architecture for the identified cross-modal learning problem and propose loss functions to opportunistically engage scene flow estimation using multiple cross-modal constraints for effective model training. Extensive experiments show the state-of-the-art performance of our method and demonstrate the effectiveness of cross-modal super-vised learning to infer more accurate 4D radar scene flow. We also show its usefulness to two subtasks - motion segmentation and ego-motion estimation. Our source code will be available on https://github.com/Toytiny/CMFlow. Fangqiang Ding, Andras Palffy, Dariu Gavrila, Xiaoxuan Lu 0001 |
CVPR | 3 |
| 2023 | Globally Guided Trajectory Planning in Dynamic EnvironmentsabstractNavigating mobile robots through environments shared with humans is challenging. From the perspective of the robot, humans are dynamic obstacles that must be avoided. These obstacles make the collision-free space nonconvex, which leads to two distinct passing behaviors per obstacle (passing left or right). For local planners, such as receding-horizon trajectory optimization, each behavior presents a local optimum in which the planner can get stuck. This may result in slow or unsafe motion even when a better plan exists. In this work, we identify trajectories for multiple locally optimal driving behaviors, by considering their topology. This identification is made consistent over successive iterations by propagating the topology information. The most suitable high-level trajectory guides a local optimization-based planner, resulting in fast and safe motion plans. We validate the proposed planner on a mobile robot in simulation and real-world experiments. Oscar de Groot, Laura Ferranti, Dariu Gavrila, Javier Alonso-Mora |
ICRA | 3 |
| 2022 | Learning to Predict Motion from Raw 3D Object DetectionsabstractWe show how to design a motion prediction algorithm that works with 3D object detections and map locations. In particular, we obtain object id’s – even though the training data does not contain any object id’s – across multiple time-steps into the future by propagating a Gaussian Mixture of likely object (e.g., vehicle) locations through time.We validate our approach on the nuScenes dataset. First, we find that a motion prediction algorithm without tracking id’s performs as well as motion prediction algorithm with tracking id’s in the training data. Second, the 3D labels of an on-board perception system are inferior (e.g., loss of detections, positional uncertainty) to those generated by offline labelling (automatic labelling pipeline, manual labelling). Even so, we find that a moderate increase in the size of the training data offsets the deterioration in prediction performance (with no additional offline labelling). Christian Neumeyer, Mario Bijelic, Dariu Gavrila |
IV | 3 |
| 2022 | Structural Knowledge Distillation for Object DetectionabstractKnowledge Distillation (KD) is a well-known training paradigm in deep neural networks where knowledge acquired by a large teacher model is transferred to a small student.KD has proven to be an effective technique to significantly improve the student's performance for various tasks including object detection. As such, KD techniques mostly rely on guidance at the intermediate feature level, which is typically implemented by minimizing an $\ell_{p}$-norm distance between teacher and student activations during training. In this paper, we propose a replacement for the pixel-wise independent $\ell_{p}$-norm based on the structural similarity (SSIM).By taking into account additional contrast and structural cues, more information within intermediate feature maps can be preserved. Extensive experiments on MSCOCO demonstrate the effectiveness of our method across different training schemes and architectures. Our method adds only little computational overhead, is straightforward to implement and at the same time it significantly outperforms the standard $\ell_p$-norms.Moreover, more complex state-of-the-art KD methods using attention-based sampling mechanisms are outperformed, including a +3.5 AP gain using a Faster R-CNN R-50 compared to a vanilla model. Philip de Rijk, Lukas Schneider, Marius Cordts, Dariu Gavrila |
NeurIPS | 4 |
| 2022 | Semantic Scene Completion Using Local Deep Implicit Functions on LiDAR DataabstractSemantic scene completion is the task of jointly estimating 3D geometry and semantics of objects and surfaces within a given extent. This is a particularly challenging task on real-world data that is sparse and occluded. We propose a scene segmentation network based on local Deep Implicit Functions as a novel learning-based method for scene completion. Unlike previous work on scene completion, our method produces a continuous scene representation that is not based on voxelization. We encode raw point clouds into a latent space locally and at multiple spatial resolutions. A global scene completion function is subsequently assembled from the localized function patches. We show that this continuous representation is suitable to encode geometric and semantic properties of extensive outdoor scenes without the need for spatial discretization (thus avoiding the trade-off between level of scene detail and the scene extent that can be covered). We train and evaluate our method on semantically annotated LiDAR scans from the Semantic KITTI dataset. Our experiments verify that our method generates a powerful representation that can be decoded into a dense 3D description of a given scene. The performance of our method surpasses the state of the art on the Semantic KITTI Scene Completion Benchmark in terms of geometric completion intersection-over-union (IoU). Christoph Rist, David Emmerichs, Markus Enzweiler, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Simple Pair Pose - Pairwise Human Pose Estimation in Dense Urban Traffic ScenesabstractDespite the success of deep learning, human pose estimation remains a challenging problem in particular in dense urban traffic scenarios. Its robustness is important for followup tasks like trajectory prediction and gesture recognition. We are interested in human pose estimation in crowded scenes with overlapping pedestrians, in particular pairwise constellations. We propose a new top-down method that relies on pairwise detections as input and jointly estimates the two poses of such pairs in a single forward pass within a deep convolutional neural network. As availability of automotive datasets providing poses and a fair amount of crowded scenes is limited, we extend the EuroCity Persons dataset by additional images and pose annotations. With 46,975 images and poses of 279,329 persons our new EuroCity Persons Dense Pose dataset is the largest pose dataset recorded from a moving vehicle. In our experiments using this dataset we show improved performance for poses of pedestrian pairs in comparison with a state of the art method for human pose estimation in crowds. Markus Braun 0003, Fabian Flohr, Sebastian Krebs, Ulrich Kreße, Dariu Gavrila |
IV | 5 |
| 2021 | Towards the detection of driver-pedestrian eye contactabstractNon-verbal communication, such as eye contact between drivers and pedestrians, has been regarded as one way to reduce accident risk. So far, studies have assumed rather than objectively measured the occurrence of eye contact. We address this research gap by developing an eye contact detection method and testing it in an indoor experiment with scripted driver–pedestrian interactions at a pedestrian crossing. Thirty participants acted as a pedestrian either standing on an imaginary curb or crossing an imaginary one-lane road in front of a stationary vehicle with an experimenter in the driver’s seat. In half of the trials, pedestrians were instructed to make eye contact with the driver; in the other half, they were prohibited from doing so. Both parties’ gaze was recorded using eye trackers. An in-vehicle stereo camera recorded the car’s point of view, a head-mounted camera recorded the pedestrian’s point of view, and the location of the driver’s and pedestrian’s eyes was estimated using image recognition. We demonstrate that eye contact can be detected by measuring the angles between the vector joining the estimated location of the driver’s and pedestrian’s eyes, and the pedestrian’s and driver’s instantaneous gaze directions, respectively, and identifying whether these angles fall below a threshold of 4°. We achieved 100% correct classification of the trials involving eye contact and those without eye contact, based on measured eye contact duration. The proposed eye contact detection method may be useful for future research into eye contact. Vishal Onkhar, Pavlo Bazilinskyy, Jork C. J. Stapel, Dimitra Dodou, Dariu Gavrila, Joost C. F. de Winter |
Pervasive Mob. Comput. | 5 |
| 2020 | ECP2.5D - Person Localization in Traffic Scenesabstract3D localization of persons from a single image is a challenging problem, where advances are largely data-driven. In this paper, we enhance the recently released EuroCity Persons detection dataset, a large and diverse automotive dataset covering pedestrians and riders. Previously, only 2D annotations and image data were provided. We introduce an automatic 3D lifting procedure by using additional LiDAR distance measurements, to augment a large part of the reasonable subset of 2D box annotations with their corresponding 3D point positions (136K persons in 46K frames of day- and night-time). The resulting dataset (coined ECP2.5D), now including Li-DAR data as well as the generated annotations, is made publicly available for (non-commercial) benchmarking of camera-based and/or LiDAR 3D object detection methods. We provide baseline results for 3D localization from single images by extending the YOLOv3 2D object detector with a distance regression including uncertainty estimation. Markus Braun 0003, Sebastian Krebs, Dariu Gavrila |
IV | 3 |
| 2020 | SCSSnet: Learning Spatially-Conditioned Scene Segmentation on LiDAR Point CloudsabstractThis work proposes a spatially-conditioned neural network to perform semantic segmentation and geometric scene completion in 3D on real-world LiDAR data. Spatially-conditioned scene segmentation (SCSSnet) is a representation suitable to encode properties of large 3D scenes at high resolution. A novel sampling strategy encodes free space information from LiDAR scans explicitly and is both simple and effective. We avoid the need for synthetically generated or volumetric ground truth data and are able to train and evaluate our method on semantically annotated LiDAR scans from the Semantic KITTI dataset. Ultimately, our method is able to predict scene geometry as well as a diverse set of semantic classes over a large spatial extent at arbitrary output resolution instead of a fixed discretization of space. Our experiments confirm that the learned scene representation is versatile and powerful and can be used for multiple downstream tasks. We perform point-wise semantic segmentation, point-of-view depth completion and ground plane segmentation. The semantic segmentation performance of our method surpasses the state of the art by a significant margin of 7% mIoU. Christoph Rist, Markus Enzweiler, Dariu Gavrila |
IV | 4 |
| 2020 | An Experimental Study on 3D Person Localization in Traffic ScenesabstractThis paper presents an experimental study on 3D person localization (i.e. pedestrians, cyclists) in traffic scenes, using monocular vision and LiDAR data. We first analyze the detection performance of two top-ranking methods (PointPillars and AVOD) on the KITTI benchmark, with respect to varying Intersection over Union (IoU) settings and the underlying parameters of 3D bounding box location, extent and orientation. Given that the KITTI dataset contains relatively few 3D person instances, we also consider the new EuroCity Persons 2.5D (ECP2.5D) dataset, which is one order of magnitude larger. We perform domain transfer experiments between the KITTI and ECP2.5D datasets, to examine how these datasets generalize with respect to each other. Joram R. Van Der Sluis, Ewoud A. I. Pool, Dariu Gavrila |
IV | 3 |
| 2019 | Privacy Protection in Street-View Panoramas Using Depth and Multi-View ImageryabstractThe current paradigm in privacy protection in street-view images is to detect and blur sensitive information. In this paper, we propose a framework that is an alternative to blurring, which automatically removes and inpaints moving objects (e.g. pedestrians, vehicles) in street-view imagery. We propose a novel moving object segmentation algorithm exploiting consistencies in depth across multiple street-view images that are later combined with the results of a segmentation network. The detected moving objects are removed and inpainted with information from other views, to obtain a realistic output image such that the moving object is not visible anymore. We evaluate our results on a dataset of 1000 images to obtain a peak noise-to-signal ratio (PSNR) and L 1 loss of 27.2 dB and 2.5%, respectively. To assess overall quality, we also report the results of a survey conducted on 35 professionals, asked to visually inspect the images whether object removal and inpainting had taken place. The inpainting dataset will be made publicly available for scientific benchmarking purposes at https://research.cyclomedia.com/. Ries Uittenbogaard, Clint Sebastian, Julien A. Vijverberg, Bas Boom, Dariu Gavrila, Peter H. N. de With |
CVPR | 5 |
| 2019 | SafeVRU: A Research Platform for the Interaction of Self-Driving Vehicles with Vulnerable Road UsersabstractThis paper presents our research platform Safe VRU for the interaction of self-driving vehicles with Vulnerable Road Users (VRUs, i.e., pedestrians and cyclists). The paper details the design (implemented with a modular structure within ROS) of the full stack of vehicle localization, environment perception, motion planning, and control, with emphasis on the environment perception and planning modules. The environment perception detects the VRUs using a stereo camera and predicts their paths with Dynamic Bayesian Networks (DBNs), which can account for switching dynamics. The motion planner is based on model predictive contouring control (MPCC) and takes into account vehicle dynamics, control objectives (e.g., desired speed), and perceived environment (i.e., the predicted VRU paths with behavioral uncertainties) over a certain time horizon. We present simulation and real-world results to illustrate the ability of our vehicle to plan and execute collision-free trajectories in the presence of VRUs. Laura Ferranti, Bruno Brito, Ewoud A. I. Pool, Ronald M. Ensing, Riender Happee, Barys N. Shyrokau, Julian F. P. Kooij, Javier Alonso-Mora, Dariu Gavrila |
IV | 10 |
| 2019 | Instance Stixels: Segmenting and Grouping Stixels into ObjectsabstractState-of-the-art stixel methods fuse dense stereo and semantic class information, e.g. from a Convolutional Neural Network (CNN), into a compact representation of driveable space, obstacles, and background. However, they do not explicitly differentiate instances within the same class. We investigate several ways to augment single-frame stixels with instance information, which can similarly be extracted by a CNN from the color input. As a result, our novel Instance Stixels method efficiently computes stixels that do account for boundaries of individual objects, and represents individual instances as grouped stixels that express connectivity. Experiments on Cityscapes demonstrate that including instance information into the stixel computation itself, rather than as a post-processing step, increases Instance AP performance with approximately the same number of stixels. Qualitative results confirm that segmentation improves, especially for overlapping objects of the same class. Additional tests with ground truth instead of CNN output show that the approach has potential for even larger gains. Our Instance Stixels software is made freely available for non-commercial research purposes. Thomas M. Hehn, Julian F. P. Kooij, Dariu Gavrila |
IV | 3 |
| 2019 | Composable Q- Functions for Pedestrian Car InteractionsabstractWe propose a novel algorithm that predicts the interaction of pedestrians with cars within a Markov Decision Process framework. It leverages the fact that Q-functions may be composed in the maximum-entropy framework, thus the solutions of two sub-tasks may be combined to approximate the full interaction problem. Sub-task one is the interaction-free navigation of a pedestrian in an urban environment and sub-task two is the interaction with an approaching car (deceleration, waiting etc.) without accounting for the environmental context (e.g. street layout). We propose a regularization scheme motivated by the soft-Bellman-equations and illustrate its necessity. We then analyze the properties of the algorithm in detail with a toy model. We find that as long as the interaction-free sub-task is modelled well with a Q-function, we can learn a representation of the interaction between a pedestrian and a car. Christian Muench, Dariu Gavrila |
IV | 2 |
| 2019 | Occlusion aware sensor fusion for early crossing pedestrian detectionabstractEarly and accurate detection of crossing pedestrians is crucial in automated driving to execute emergency manoeuvres in time. This is a challenging task in urban scenarios however, where people are often occluded (not visible) behind objects, e.g. other parked vehicles. In this paper, an occlusion aware multi-modal sensor fusion system is proposed to address scenarios with crossing pedestrians behind parked vehicles. Our proposed method adjusts the detection rate in different areas based on sensor visibility. We argue that using this occlusion information can help to evaluate the measurements. Our experiments on real world data show that fusing radar and stereo camera for such tasks is beneficial, and that including occlusion into the model helps to detect pedestrians earlier and more accurately. Andras Palffy, Julian F. P. Kooij, Dariu Gavrila |
IV | 3 |
| 2019 | Context-based cyclist path prediction using Recurrent Neural NetworksabstractThis paper proposes a Recurrent Neural Network (RNN) for cyclist path prediction to learn the effect of contextual cues on the behavior directly in an end- to-end approach, removing the need for any annotations. The proposed RNN incorporates three distinct contextual cues: one related to actions of the cyclist, one related to the location of the cyclist on the road, and one related to the interaction between the cyclist and the egovehicle. The RNN predicts a Gaussian distribution over the future position of the cyclist one second into the future with a higher accuracy, compared to a current state-of-the-art model that is based on dynamic mode annotations, where our model attains an average prediction error of 33 cm one second into the future. Ewoud A. I. Pool, Julian F. P. Kooij, Dariu Gavrila |
IV | 3 |
| 2019 | Cross-Sensor Deep Domain Adaptation for LiDAR Detection and SegmentationabstractA considerable amount of annotated training data is necessary to achieve state-of-the-art performance in perception tasks using point clouds. Unlike RGB-images, LiDAR point clouds captured with different sensors or varied mounting positions exhibit a significant shift in their input data distribution. This can impede transfer of trained feature extractors between datasets as it degrades performance vastly. We analyze the transferability of point cloud features between two different LiDAR sensor set-ups (32 and 64 vertical scanning planes with different geometry). We propose a supervised training methodology to learn transferable features in a pre-training step on LiDAR datasets that are heterogeneous in their data and label domains. In extensive experiments on object detection and semantic segmentation in a multi-task setup we analyze the performance of our network architecture under the impact of a change in the input data domain. We show that our pre-training approach effectively increases performance for both target tasks at once without having an actual multi-task dataset available for pre-training. Christoph Rist, Markus Enzweiler, Dariu Gavrila |
IV | 3 |
| 2019 | DD-Pose - A large-scale Driver Head Pose BenchmarkabstractWe introduce DD-Pose, the Daimler TU Delft Driver Head Pose Benchmark, a large-scale and diverse benchmark for image-based head pose estimation and driver analysis. It contains 330k measurements from multiple cameras acquired by an in-car setup during naturalistic drives. Large out-of-plane head rotations and occlusions are induced by complex driving scenarios, such as parking and driver-pedestrian interactions. Precise head pose annotations are obtained by a motion capture sensor and a novel calibration device. A high resolution stereo driver camera is supplemented by a camera capturing the driver cabin. Together with steering wheel and vehicle motion information, DD-Pose paves the way for holistic driver analysis. Our experiments show that the new dataset offers a broad distribution of head poses, comprising an order of magnitude more samples of rare poses than a comparable dataset. By an analysis of a state-of-the-art head pose estimation method, we demonstrate the challenges offered by the benchmark. The dataset and evaluation code are made freely available to academic and non-profit institutions for non-commercial benchmarking purposes. Markus Roth, Dariu Gavrila |
IV | 2 |
| 2019 | Context-Based Path Prediction for Targets with Switching DynamicsabstractAnticipating future situations from streaming sensor data is a key perception challenge for mobile robotics and automated vehicles. We address the problem of predicting the path of objects with multiple dynamic modes. The dynamics of such targets can be described by a Switching Linear Dynamical System (SLDS). However, predictions from this probabilistic model cannot anticipate when a change in dynamic mode will occur. We propose to extract various types of cues with computer vision to provide context on the target’s behavior, and incorporate these in a Dynamic Bayesian Network (DBN). The DBN extends the SLDS by conditioning the mode transition probabilities on additional context states. We describe efficient online inference in this DBN for probabilistic path prediction, accounting for uncertainty in both measurements and target behavior. Our approach is illustrated on two scenarios in the Intelligent Vehicles domain concerning pedestrians and cyclists, so-called Vulnerable Road Users (VRUs). Here, context cues include the static environment of the VRU, its dynamic environment, and its observed actions. Experiments using stereo vision data from a moving vehicle demonstrate that the proposed approach results in more accurate path prediction than SLDS at the relevant short time horizon (1 s). It slightly outperforms a computationally more demanding state-of-the-art method. Julian F. P. Kooij, Fabian Flohr, Ewoud A. I. Pool, Dariu Gavrila |
Int. J. Comput. Vis. | 4 |
| 2019 | EuroCity Persons: A Novel Benchmark for Person Detection in Traffic ScenesabstractBig data has had a great share in the success of deep learning in computer vision. Recent works suggest that there is significant further potential to increase object detection performance by utilizing even bigger datasets. In this paper, we introduce the EuroCity Persons dataset, which provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes. The images for this dataset were collected on-board a moving vehicle in 31 cities of 12 European countries. With over 238200 person instances manually labeled in over 47300 images, EuroCity Persons is nearly one order of magnitude larger than datasets used previously for person detection in traffic scenes. The dataset furthermore contains a large number of person orientation annotations (over 211200). We optimize four state-of-the-art deep learning approaches (Faster R-CNN, R-FCN, SSD and YOLOv3) to serve as baselines for the new object detection benchmark. We analyze the generalization capabilities of these detectors when trained with the new dataset. We furthermore study the effect of the training set size, the dataset diversity (day- vs. night-time, geographical region), the dataset detail (i.e. availability of object orientation information) and the annotation quality on the detector performance. Finally, we analyze error sources and discuss the road ahead. Markus Braun 0003, Sebastian Krebs, Fabian Flohr, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Using road topology to improve cyclist path predictionabstractWe learn motion models for cyclist path prediction on real-world tracks obtained from a moving vehicle, and propose to exploit the local road topology to obtain better predictive distributions. The tracks are extracted from the Tsinghua-Daimler Cyclist Benchmark for cyclist detection, and corrected for vehicle egomotion. Tracks are then spatially aligned to local curves and crossings in the road. We study a standard approach for path prediction in the literature based on Kalman Filters, as well as a mixture of specialized filters related to specific road orientations at junctions. Our experiments demonstrate an improved prediction accuracy (up to 20% on sharp turns) of mixing specialized motion models for canonical directions, and prior knowledge on the road topology. The new track data complements the existing video, disparity and annotation data of the original benchmark, and will be made publicly available. Ewoud A. I. Pool, Julian F. P. Kooij, Dariu Gavrila |
Intelligent Vehicles Symposium | 3 |
| 2017 | A Unified Framework for Concurrent Pedestrian and Cyclist DetectionabstractExtensive research interest has been focused on protecting vulnerable road users in recent years, particularly pedestrians and cyclists, due to their attributes of vulnerability. However, comparatively little effort has been spent on detecting pedestrian and cyclist together, particularly when it concerns quantitative performance analysis on large datasets. In this paper, we present a unified framework for concurrent pedestrian and cyclist detection, which includes a novel detection proposal method (termed UB-MPR) to output a set of object candidates, a discriminative deep model based on Fast R-CNN for classification and localization, and a specific postprocessing step to further improve detection performance. Experiments are performed on a new pedestrian and cyclist dataset containing 30 490 annotated pedestrian and 26 771 cyclist instances in over 50 000 images, recorded from a moving vehicle in the urban traffic of Beijing. Experimental results indicate that the proposed method outperforms other state-of-the-art methods significantly. Lingxi Li 0001, Fabian Flohr, Jianqiang Wang 0003, Hui Xiong 0006, Bernhard Morys, Shuyue Pan, Dariu Gavrila, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2016 | A new benchmark for vision-based cyclist detectionabstractSignificant progress has been achieved over the past decade on vision-based pedestrian detection; this has led to active pedestrian safety systems being deployed in most mid- to high-range cars on the market. Comparatively little effort has been spent on vision-based cyclist detection, especially when it concerns quantitative performance analysis on large datasets. We present a large-scale experimental study on cyclist detection where we examine the currently most promising object detection methods; we consider Aggregated Channel Features, Deformable Part Models and Region-based Convolutional Neural Networks. We also introduce a new method called Stereo-Proposal based Fast R-CNN (SP-FRCN) to detect cyclists based on stereo proposals and Fast R-CNN (FRCN) framework. Experiments are performed on a dataset containing 22161 annotated cyclist instances in over 30000 images, recorded from a moving vehicle in the urban traffic of Beijing. Results indicate that all the three solution families can reach top performance around 0.89 average precision on the easy case, but the performance drops gradually with the difficulty increasing. The dataset including rich annotations, stereo images and evaluation scripts (termed “Tsinghua-Daimler Cyclist Benchmark”) is made public to the scientific community, to serve as a common point of reference for future research. Fabian Flohr, Hui Xiong 0006, Markus Braun 0003, Shuyue Pan, Keqiang Li 0002, Dariu Gavrila |
Intelligent Vehicles Symposium | 8 |
| 2016 | Driver and pedestrian awareness-based collision risk analysisabstractWe present a novel approach for vehicle-pedestrian collision risk analysis that incorporates mutual situational awareness, a degree of potential motion coupling and the spatial layout of the environment. The approach uses a Dynamic Bayesian Network (DBN) for modeling the individual object paths; collision risk is subsequently computed by an intersection operation. More specifically, the proposed DBN consists of two subgraphs for modeling pedestrian and vehicle path, respectively. They consist of latent states on top of Switching Linear Dynamical Systems (SLDSs) to anticipate changes in object dynamics. The pedestrian and vehicle-related sub-graphs contain latent states to model whether the pedestrian has seen the oncoming vehicle, and conversely, whether the driver has seen the pedestrian (associated measurements involve the respective head orientations). The pedestrian-related sub-graph furthermore contains a latent state modeling whether the pedestrian is at the curbside or not. Finally, a latent state is shared by the two sub-graphs, which models the potential motion coupling (i.e. at full awareness of the other traffic participant). We consider the scenario of a crossing pedestrian, who might stop or continue walking at the curb, in combination with an approaching vehicle, that might stop or continue driving. In experiments we illustrate that with the proposed approach, a more anticipatory driver warning and/or vehicle control strategy can be implemented. Markus Roth, Fabian Flohr, Dariu Gavrila |
Intelligent Vehicles Symposium | 3 |
| 2016 | Multi-modal human aggression detectionabstractThis paper presents a smart surveillance system named CASSANDRA, aimed at detecting instances of aggressive human behavior in public environments. A distinguishing aspect of CASSANDRA is the exploitation of complementary audio and video cues to disambiguate scene activity in real-life environments. From the video side, the system uses overlapping cameras to track persons in 3D and to extract features regarding the limb motion relative to the torso. From the audio side, it classifies instances of speech, screaming, singing, and kicking-object. The audio and video cues are fused with contextual cues (interaction, auxiliary objects); a Dynamic Bayesian Network (DBN) produces an estimate of the ambient aggression level. Our prototype system is validated on a realistic set of scenarios performed by professional actors at an actual train station to ensure a realistic audio and video noise setting. Julian F. P. Kooij, Martijn Liem, Johannes D. Krijnders, Tjeerd C. Andringa, Dariu Gavrila |
Comput. Vis. Image Underst. | 5 |
| 2016 | Mixture of Switching Linear Dynamics to Discover Behavior Patterns in Object TracksabstractWe present a novel non-parametric Bayesian model to jointly discover the dynamics of low-level actions and high-level behaviors of tracked objects. In our approach, actions capture both linear, low-level object dynamics, and an additional spatial distribution on where the dynamic occurs. Furthermore, behavior classes capture high-level temporal motion dependencies in Markov chains of actions, thus each learned behavior is a switching linear dynamical system. The number of actions and behaviors is discovered from the data itself using Dirichlet Processes. We are especially interested in cases where tracks can exhibit large kinematic and spatial variations, e.g. person tracks in open environments, as found in the visual surveillance and intelligent vehicle domains. The model handles real-valued features directly, so no information is lost by quantizing measurements into 'visual words', and variations in standing, walking and running can be discovered without discrete thresholds. We describe inference using Markov Chain Monte Carlo sampling and validate our approach on several artificial and real-world pedestrian track datasets from the surveillance and intelligent vehicle domain. We show that our model can distinguish between relevant behavior patterns that an existing state-of-the-art hierarchical model for clustering and simpler model variants cannot. The software and the artificial and surveillance datasets are made publicly available for benchmarking purposes. Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Identifying multiple objects from their appearance in inaccurate detections
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
Comput. Vis. Image Underst. | 3 |
| 2015 | A Probabilistic Framework for Joint Pedestrian Head and Body Orientation EstimationabstractWe present a probabilistic framework for the joint estimation of pedestrian head and body orientation from a mobile stereo vision platform. For both head and body parts, we convert the responses of a set of orientation-specific detectors into a (continuous) probability density function. The parts are localized by means of apictorial structureapproach, which balances part-based detector responses with spatial constraints. Head and body orientations are estimated jointly to account for anatomical constraints. The joint single-frame orientation estimates are integrated over time by particle filtering. The experiments involved data from a vehicle-mounted stereo vision camera in a realistic traffic setting; 65 pedestrian tracks were supplied by a state-of-the-art pedestrian tracker. We show that the proposed joint probabilistic orientation estimation framework reduces the mean absolute head and body orientation error up to 15° compared with simpler methods. This results in a mean absolute head/body orientation error of about 21°/19°, which remains fairly constant up to a distance of 25 m. Our system currently runs in near real time (8–9 Hz). Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2014 | Context-Based Pedestrian Path Prediction
Julian F. P. Kooij, Nicolas Schneider, Fabian Flohr, Dariu Gavrila |
ECCV (6) | 4 |
| 2014 | The development and real-world deployment of FROG, the fun robotic outdoor guideabstractThis video details the development of an intelligent outdoor Guide robot. The main objective is to deploy an innovative robotic guide which is not only able to show information, but to react to the affective states of the users, and to offer location-based services using augmented reality. The scientific challenges concern autonomous outdoor navigation and localization, robust 24/7 operation, affective interaction with visitors through outdoor human and facial feature detection as well as engaging interactive behaviors in an ongoing non-verbal dialogue with the user. Vanessa Evers, Nuno Menezes, Luis Merino, Dariu Gavrila, Fernando Nabais, Maja Pantic, Paulo Alvito, Daphne E. Karreman |
HRI | 4 |
| 2014 | Joint probabilistic pedestrian head and body orientation estimationabstractWe present an approach for the joint probabilistic estimation of pedestrian head and body orientation in the context of intelligent vehicles. For both, head and body, we convert the output of a set of orientation-specific detectors into a full (continuous) probability density function. The parts are localized with a pictorial structure approach which balances part-based detector output with spatial constraints. Head and body orientation estimates are furthermore coupled probabilistically to account for anatomical constraints. Finally, the coupled single-frame orientation estimates are integrated over time by particle filtering. The experiments involve 37 pedestrian tracks obtained from an external stereo vision-based pedestrian detector in realistic traffic settings. We show that the proposed joint probabilistic orientation estimation approach reduces the mean head and body orientation error by 10 degrees and more. Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila |
Intelligent Vehicles Symposium | 4 |
| 2014 | Analysis of pedestrian dynamics from a vehicle perspectiveabstractAccurate motion models are key to many tasks in the intelligent vehicle domain, but simple Linear Dynamics (e.g. Kalman filtering) do not exploit the spatio-temporal context of motion. We present a method to learn Switching Linear Dynamics of object tracks observed from within a driving vehicle. Each switching state captures object dynamics as a mean motion with variance, but also has an additional spatial distribution on where the dynamic is seen relative to the vehicle. Thus, both an object's previous movements and current location will make certain dynamics more probable for subsequent time steps. To train the model, we use Bayesian inference to sample parameters from the posterior, and jointly learn the required number of dynamics. Unlike Maximum Likelihood learning, inference is robust against overfitting and poor initialization. We demonstrate our approach on an ego-motion compensated track dataset of pedestrians, and illustrate how the switching dynamics can make more accurate path predictions than a mixture of linear dynamics for crossing pedestrians. Julian F. P. Kooij, Nicolas Schneider, Dariu Gavrila |
Intelligent Vehicles Symposium | 3 |
| 2014 | Joint multi-person detection and tracking from overlapping cameras
Martijn Liem, Dariu Gavrila |
Comput. Vis. Image Underst. | 2 |
| 2014 | Coupled person orientation estimation and appearance modeling using spherical harmonics
Martijn Liem, Dariu Gavrila |
Image Vis. Comput. | 2 |
| 2013 | PedCut: an iterative framework for pedestrian segmentation combining shape models and multiple data cuesabstractThis paper presents an iterative, EM-like framework for accurate pedestrian segmentation, combining generative shape models and multiple data cues.In the E-step, shape priors are introduced in the unary terms of a Conditional Random Field (CRF) formulation, joining other data terms derived from color, texture and disparity cues.In the M-step, the resulting segmentation is used to adapt an Active Shape Model (ASM), after which the EM process alternates.Experiments on the public Penn-Fudan pedestrian dataset suggest that our method outperforms the state-of-the-art.We further provide results on a new Daimler pedestrian dataset, captured from on-board a vehicle, which includes disparity data.This dataset is made public to facilitate benchmarking. Fabian Flohr, Dariu Gavrila |
BMVC | 2 |
| 2013 | A Comparative Study on Multi-person Tracking Using Overlapping Cameras
Martijn Liem, Dariu Gavrila |
ICVS | 2 |
| 2012 | A Non-parametric Hierarchical Model to Discover Behavior Dynamics from Tracks
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila |
ECCV (6) | 3 |
| 2012 | Multi-view 3D Human Pose Estimation in Complex EnvironmentabstractWe introduce a framework for unconstrained 3D human upper body pose estimation from multiple camera views in complex environment. Its main novelty lies in the integration of three components: single-frame pose recovery, temporal integration and model texture adaptation. Single-frame pose recovery consists of a hypothesis generation stage, in which candidate 3D poses are generated, based on probabilistic hierarchical shape matching in each camera view. In the subsequent hypothesis verification stage, the candidate 3D poses are re-projected into the other camera views and ranked according to a multi-view likelihood measure. Temporal integration consists of computing K-best trajectories combining a motion model and observations in a Viterbi-style maximum-likelihood approach. Poses that lie on the best trajectories are used to generate and adapt a texture model, which in turn enriches the shape likelihood measure used for pose recovery. The multiple trajectory hypotheses are used to generate pose predictions, augmenting the 3D pose candidates generated at the next time step. We demonstrate that our approach outperforms the state-of-the-art in experiments with large and challenging real-world data from an outdoor setting. Michael Hofmann 0005, Dariu Gavrila |
Int. J. Comput. Vis. | 2 |
| 2011 | A new benchmark for stereo-based pedestrian detectionabstractPedestrian detection is a rapidly evolving area in the intelligent vehicles domain. Stereo vision is an attractive sensor for this purpose. But unlike for monocular vision, there are no realistic, large scale benchmarks available for stereo-based pedestrian detection, to provide a common point of reference for evaluation. This paper introduces the Daimler Stereo-Vision Pedestrian Detection benchmark, which consists of several thousands of pedestrians in the training set, and a 27-min test drive through urban environment and associated vehicle data. The data, including ground truth, is made publicly available for non-commercial purposes. The paper furthermore quantifies the benefit of stereo vision for ROI generation and localization; at equal detection rates, false positives are reduced by a factor of 4-5 with stereo over mono, using the same HOG/linSVM classification component. Christoph Gustav Keller, Markus Enzweiler, Dariu Gavrila |
Intelligent Vehicles Symposium | 3 |
| 2011 | 3D Human model adaptation by frame selection and shape-texture optimization
Michael Hofmann 0005, Dariu Gavrila |
Comput. Vis. Image Underst. | 2 |
| 2011 | A Multilevel Mixture-of-Experts Framework for Pedestrian ClassificationabstractNotwithstanding many years of progress, pedestrian recognition is still a difficult but important problem. We present a novel multilevel Mixture-of-Experts approach to combine information from multiple features and cues with the objective of improved pedestrian classification. On pose-level, shape cues based on Chamfer shape matching provide sample-dependent priors for a certain pedestrian view. On modality-level, we represent each data sample in terms of image intensity, (dense) depth, and (dense) flow. On feature-level, we consider histograms of oriented gradients (HOG) and local binary patterns (LBP). Multilayer perceptrons (MLP) and linear support vector machines (linSVM) are used as expert classifiers. Experiments are performed on a unique real-world multi-modality dataset captured from a moving vehicle in urban traffic. This dataset has been made public for research purposes. Our results show a significant performance boost of up to a factor of 42 in reduction of false positives at constant detection rates of our approach compared to a baseline intensity-only HOG/linSVM approach. Markus Enzweiler, Dariu Gavrila |
IEEE Trans. Image Process. | 2 |
| 2011 | Active Pedestrian Safety by Automatic Braking and Evasive SteeringabstractActive safety systems hold great potential for reducing accident frequency and severity by warning the driver and/or exerting automatic vehicle control ahead of crashes. This paper presents a novel active pedestrian safety system that combines sensing, situation analysis, decision making, and vehicle control. The sensing component is based on stereo vision, and it fuses the following two complementary approaches for added robustness: 1) motion-based object detection and 2) pedestrian recognition. The highlight of the system is its ability to decide, within a split second, whether it will perform automatic braking or evasive steering and reliably execute this maneuver at relatively high vehicle speed (up to 50 km/h). We performed extensive precrash experiments with the system on the test track (22 scenarios with real pedestrians and a dummy). We obtained a significant benefit in detection performance and improved lateral velocity estimation by the fusion of motion-based object detection and pedestrian recognition. On a fully reproducible scenario subset, involving the dummy that laterally enters into the vehicle path from behind an occlusion, the system executed, in more than 40 trials, the intended vehicle action, i.e., automatic braking (if a full stop is still possible) or automatic evasive steering. Christoph Gustav Keller, Thao Dang 0002, Hans Fritz, Armin Joos, Clemens Rabe, Dariu Gavrila |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2011 | The Benefits of Dense Stereo for Pedestrian DetectionabstractThis paper presents a novel pedestrian detection system for intelligent vehicles. We propose the use of dense stereo for both the generation of regions of interest and pedestrian classification. Dense stereo allows the dynamic estimation of camera parameters and the road profile, which, in turn, provides strong scene constraints on possible pedestrian locations. For classification, we extract spatial features (gradient orientation histograms) directly from dense depth and intensity images. Both modalities are represented in terms of individual feature spaces, in which discriminative classifiers (linear support vector machines) are learned. We refrain from the construction of a joint feature space but instead employ a fusion of depth and intensity on the classifier level. Our experiments involve challenging image data captured in complex urban environments (i.e., undulating roads and speed bumps). Our results show a performance improvement by up to a factor of 7.5 at the classification level and up to a factor of 5 at the tracking level (reduction in false alarms at constant detection rates) over a system with static scene constraints and intensity-only classification. Christoph Gustav Keller, Markus Enzweiler, Marcus Rohrbach, David Fernández Llorca, Christoph Schnörr, Dariu Gavrila |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2010 | Multi-cue pedestrian classification with partial occlusion handlingabstractThis paper presents a novel mixture-of-experts framework for pedestrian classification with partial occlusion handling. The framework involves a set of component-based expert classifiers trained on features derived from intensity, depth and motion. To handle partial occlusion, we compute expert weights that are related to the degree of visibility of the associated component. This degree of visibility is determined by examining occlusion boundaries, i.e. discontinuities in depth and motion. Occlusion-dependent component weights allow to focus the combined decision of the mixture-of-experts classifier on the unoccluded body parts. In experiments on extensive real-world data sets, with both partially occluded and non-occluded pedestrians, we obtain significant performance boosts over state-of-the-art approaches by up to a factor of four in reduction of false positives at constant detection rates. The dataset is made public for benchmarking purposes. Markus Enzweiler, Angela Eigenstetter, Bernt Schiele, Dariu Gavrila |
CVPR | 4 |
| 2010 | Integrated pedestrian classification and orientation estimationabstractThis paper presents a novel approach to single-frame pedestrian classification and orientation estimation. Unlike previous work which addressed classification and orientation separately with different models, our method involves a probabilistic framework to approach both in a unified fashion. We address both problems in terms of a set of view-related models which couple discriminative expert classifiers with sample-dependent priors, facilitating easy integration of other cues (e.g. motion, shape) in a Bayesian fashion. This mixture-of-experts formulation approximates the probability density of pedestrian orientation and scales-up to the use of multiple cameras. Experiments on large real-world data show a significant performance improvement in both pedestrian classification and orientation estimation of up to 50%, compared to state-of-the-art, using identical data and evaluation techniques. Markus Enzweiler, Dariu Gavrila |
CVPR | 2 |
| 2009 | Multi-person Tracking with Overlapping Cameras in Complex, Dynamic EnvironmentsabstractThis paper presents a multi-camera system to track multiple persons in complex, dynamic environments. Position measurements are obtained by carving out the space defined by foreground regions in the overlapping camera views and projecting these onto blobs on the ground plane. Person appearance is described in terms of the colour histograms in the various camera views of three vertical body regions (head-shoulder, torso, legs). The assignment of measurements to tracks (modelled by Kalman filters) is done in a non-greedy, global fashion based on ground plane position and colour appearance. The advantage of the proposed approach is that the decision on correspondences across cameras is delayed until it can be performed at the object-level, where it is more robust. We demonstrate the effectiveness of the proposed approach using data from three cameras overlooking a complex outdoor setting (train platform), containing a significant amount of lighting and background changes. Martijn Liem, Dariu Gavrila |
BMVC | 2 |
| 2009 | Multi-view 3D human pose estimation combining single-frame recovery, temporal integration and model adaptationabstractWe present a system for the estimation of unconstrained 3D human upper body movement from multiple cameras. Its main novelty lies in the integration of three components: single frame pose recovery, temporal integration and model adaptation. Single frame pose recovery consists of a hypothesis generation stage, where candidate 3D poses are generated based on hierarchical shape matching in the individual camera views. In the subsequent hypothesis verification stage, candidate 3D poses are reprojected to the other camera views and ranked according to a multiview matching score. Temporal integration consists of computing best trajectories combining a motion model and observations in a Viterbi style maximum likelihood approach. Poses that lie on the best trajectories are used to generate and adapt a texture model, which in turn enriches the shape component used for pose recovery. We demonstrate that our approach outperforms the state of the art in experiments with large and challenging real world data from an outdoor setting. The new data set is made public to facilitate benchmarking. Michael Hofmann 0005, Dariu Gavrila |
CVPR | 2 |
| 2009 | Monocular Pedestrian Detection: Survey and ExperimentsabstractPedestrian detection is a rapidly evolving area in computer vision with key applications in intelligent vehicles, surveillance, and advanced robotics. The objective of this paper is to provide an overview of the current state of the art from both methodological and experimental perspectives. The first part of the paper consists of a survey. We cover the main components of a pedestrian detection system and the underlying models. The second (and larger) part of the paper contains a corresponding experimental study. We consider a diverse set of state-of-the-art systems: wavelet-based AdaBoost cascade [74], HOG/linSVM [11], NN/LRF [75], and combined shape-texture detection [23]. Experiments are performed on an extensive data set captured onboard a vehicle driving through urban environment. The data set includes many thousands of training samples as well as a 27-minute test sequence involving more than 20,000 images with annotated pedestrian locations. We consider a generic evaluation setting and one specific to pedestrian detection onboard a vehicle. Results indicate a clear advantage of HOG/linSVM at higher image resolutions and lower processing speeds, and a superiority of the wavelet-based AdaBoost cascade approach at lower image resolutions and (near) real-time processing speeds. The data set (8.5 GB) is made public for benchmarking purposes. Markus Enzweiler, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | A mixed generative-discriminative framework for pedestrian classificationabstractThis paper presents a novel approach to pedestrian classification which involves utilizing the synthesized virtual samples of a learned generative model to enhance the classification performance of a discriminative model. Our generative model captures prior knowledge about the pedestrian class in terms of a number of probabilistic shape and texture models, each attuned to a particular pedestrian pose. Active learning provides the link between the generative and discriminative model, in the sense that the former is selectively sampled such that the training process is guided towards the most informative samples of the latter. In large-scale experiments on real-world datasets of tens of thousands of samples, we demonstrate a significant improvement in classification performance of the combined generative-discriminative approach over the discriminative-only approach (the latter exemplified by a neural network with local receptive fields and a support vector machine using Haar wavelet features). Markus Enzweiler, Dariu Gavrila |
CVPR | 2 |
| 2008 | Pedestrian Detection and Tracking Using a Mixture of View-Based Shape-Texture ModelsabstractThis paper presents a robust multicue approach to the integrated detection and tracking of pedestrians in a cluttered urban environment. A novel spatiotemporal object representation is proposed, which combines a generative shape model and a discriminative texture classifier, both of which are composed of a mixture of pose-specific submodels. Shape is represented by a set of linear subspace models, which is an extension of point distribution models, with shape transitions being modeled by a first-order Markov process. Texture, i.e., the shape-normalized intensity pattern, is represented by a manifold that is implicitly delimited by a set of pattern classifiers, whereas texture transition is modeled by a random walk. Direct 3-D measurements that are provided by a stereo system are further incorporated into the observation density function. We employ a Bayesian framework based on particle filtering to achieve integrated object detection and tracking. Large-scale experiments that involve pedestrian detection and tracking from a moving vehicle demonstrate the benefit of the proposed approach. Stefan Munder, Christoph Schnörr, Dariu Gavrila |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2007 | Looking at peopleabstractThe ability to recognize humans and their activities by vision is key for a machine to interact intelligently and effortlessly with a human-inhabited environment. Because of many important applications, "Looking at People" is currently the most active application domain in computer vision. This talk provides an overview of recent developments, with focus on methodologies for detecting, tracking people and recognizing their activities. Dariu Gavrila |
AVSS | 1 |
| 2007 | CASSANDRA: audio-video sensor fusion for aggression detectionabstractThis paper presents a smart surveillance system named CASSANDRA, aimed at detecting instances of aggressive human behavior in public environments. A distinguishing aspect of CASSANDRA is the exploitation of the complimentary nature of audio and video sensing to disambiguate scene activity in real-life, noisy and dynamic environments. At the lower level, independent analysis of the audio and video streams yields intermediate descriptors of a scene like: "scream", "passing train" or "articulation energy". At the higher level, a Dynamic Bayesian Network is used as a fusion mechanism that produces an aggregate aggression indication for the current scene. Our prototype system is validated on a set of scenarios performed by professional actors at an actual train station to ensure a realistic audio and video noise setting. Wojtek Zajdel, Johannes D. Krijnders, Tjeerd C. Andringa, Dariu Gavrila |
AVSS | 4 |
| 2007 | Multi-cue Pedestrian Detection and Tracking from a Moving Vehicle
Dariu Gavrila, Stefan Munder |
Int. J. Comput. Vis. | 1 |
| 2007 | A Bayesian, Exemplar-Based Approach to Hierarchical Shape MatchingabstractThis paper presents a novel probabilistic approach to hierarchical, exemplar-based shape matching. No feature correspondence is needed among exemplars, just a suitable pairwise similarity measure. The approach uses a template tree to efficiently represent and match the variety of shape exemplars. The tree is generated offline by a bottom-up clustering approach using stochastic optimization. Online matching involves a simultaneous coarse-to-fine approach over the template tree and over the transformation parameters. The main contribution of this paper is a Bayesian model to estimate the a posteriori probability of the object class, after a certain match at a node of the tree. This model takes into account object scale and saliency and allows for a principled setting of the matching thresholds such that unpromising paths in the tree traversal process are eliminated early on. The proposed approach was tested in a variety of application domains. Here, results are presented on one of the more challenging domains: real-time pedestrian detection from a moving vehicle. A significant speed-up is obtained when comparing the proposed probabilistic matching approach with a manually tuned nonprobabilistic variant, both utilizing the same template tree structure. Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2006 | An Experimental Study on Pedestrian ClassificationabstractDetecting people in images is key for several important application domains in computer vision. This paper presents an in-depth experimental study on pedestrian classification; multiple feature-classifier combinations are examined with respect to their ROC performance and efficiency. We investigate global versus local and adaptive versus nonadaptive features, as exemplified by PCA coefficients, Haar wavelets, and local receptive fields (LRFs). In terms of classifiers, we consider the popular Support Vector Machines (SVMs), feed-forward neural networks, and k-nearest neighbor classifier. Experiments are performed on a large data set consisting of 4,000 pedestrian and more than 25,000 nonpedestrian (labeled) images captured in outdoor urban environments. Statistically meaningful results are obtained by analyzing performance variances caused by varying training and test sets. Furthermore, we investigate how classification performance and training sample size are correlated. Sample size is adjusted by increasing the number of manually labeled training data or by employing automatic bootstrapping or cascade techniques. Our experiments show that the novel combination of SVMs with LRF features performs best. A boosted cascade of Haar wavelets can, however, reach quite competitive results, at a fraction of computational cost. The data set used in this paper is made public, establishing a benchmark for this important problem. Stefan Munder, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Real-time dense stereo for intelligent vehiclesabstractStereo vision is an attractive passive sensing technique for obtaining three-dimensional (3-D) measurements. Recent hardware advances have given rise to a new class of real-time dense disparity estimation algorithms. This paper examines their suitability for intelligent vehicle (IV) applications. In order to gain a better understanding of the performance and the computational-cost tradeoff, the authors created a framework of real-time implementations. This consists of different methodical components based on single instruction multiple data (SIMD) techniques. Furthermore, the resulting algorithmic variations are compared with other publicly available algorithms. The authors argue that existing publicly available stereo data sets are not very suitable for the IV domain. Therefore, the authors' evaluation of stereo algorithms is based on novel realistically looking simulated data as well as real data from complex urban traffic scenes. In order to facilitate future benchmarks, all data used in this paper is made publicly available. The results from this study reveal that there is a considerable influence of scene conditions on the performance of all tested algorithms. Approaches that aim for (global) search optimization are more affected by this than other approaches. The best overall performance is achieved by the proposed multiple-window algorithm, which uses local matching and a left-right check for a robust error rejection. Timing results show that the simplest of the proposed SIMD variants are more than twice as fast than the most complex one. Nevertheless, the latter still achieves real-time processing speeds, while their average accuracy is at least equal to that of publicly available non-SIMD algorithms Wannes van der Mark, Dariu Gavrila |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2004 | A Bayesian Framework for Multi-cue 3D Object Tracking
Jan Giebel, Dariu Gavrila, Christoph Schnörr |
ECCV (4) | 2 |
| 2002 | FPGA-Based Template Matching Using Distance TransformsabstractThis paper presents a high-performance FPGA solution to generic shape-based object detection in images. The underlying detection method involves representing the target object by binary templates containing positional and directional edge information. A particular scene image is preprocessed by edge segmentation, edge cleaning and distance transforms. Matching involves correlating the templates with the distance-transformed scene image and determining the locations where the mismatch is below a certain user-defined threshold. Although successful in the past, a significant drawback of these matching methods has been their large computational cost when implemented on a sequential general-purpose processor. In this paper we present a step by step implementation of the components of such object detection systems, taking advantage of the data and logical parallelism opportunities offered by an FPGA architecture. The realization of a pipelined calculation of the preprocessing and correlation on FPGA is presented in detail. Stefan Hezel, Andreas Kugel, Reinhard Männer, Dariu Gavrila |
FCCM | 4 |
| 2001 | Virtual Sample Generation For Template-Based Shape MatchingabstractThis paper presents a method for improving the performance of matching systems that correlate using shape templates. The basic idea involves extending an existing set of training shapes with generated "virtual" shapes, in order to improve representational capability, yet no a-priori feature correspondence is necessary among the original shapes in the training set. Instead, an integrated clustering and registration approach partitions the original shape samples into clusters of similar and registered shapes; in each cluster a separate feature space is embedded. This allows the derivation of standard compact parameterizations for each cluster. This paper demonstrates that sampling these low-order spaces can produce an extended training set which facilitates a superior matching performance, as measured by a ROC curve. In the experiments, we consider a realistic application involving thousands of pedestrian shapes and perform correlation matching based on distance transforms. Dariu Gavrila, Jan Giebel |
CVPR (1) | 1 |
| 2000 | Pedestrian Detection from a Moving Vehicle
Dariu Gavrila |
ECCV (2) | 1 |
| 1999 | Real-Time Object Detection for "Smart" VehiclesabstractThis paper presents an efficient shape-based object detection method based on Distance Transforms and describes its use for real-time vision on-board vehicles. The method uses a template hierarchy to capture the variety of object shapes; efficient hierarchies can be generated offline for given shape distributions using stochastic optimization techniques (i.e. simulated annealing). Online, matching involves a simultaneous coarse-to-fine approach over the shape hierarchy and over the transformation parameters. Very large speed-up factors are typically obtained when comparing this approach with the equivalent brute-force formulation; we have measured gains of several orders of magnitudes. We present experimental results on the real-time detection of traffic signs and pedestrians from a moving vehicle. Because of the highly time sensitive nature of these vision tasks, we also discuss some hardware-specific implementations of the proposed method as far as SIMD parallelism is concerned. Dariu Gavrila, Vasanth Philomin |
ICCV | 1 |
| 1999 | The Visual Analysis of Human Movement: A Survey
Dariu Gavrila |
Comput. Vis. Image Underst. | 1 |
| 1998 | Multi-feature hierarchical template matching using distance transformsabstractWe describe a multi-feature hierarchical algorithm to efficiently match N objects (templates) with am image using distance transforms (DTs). The matching is under translation, but it can cover more general transformations by generating the various transformed templates explicitly. The novel part of the algorithm is that, in addition to a coarse-to-fine search over the translation parameters, the N templates are grouped off-line into a template hierarchy based on their similarity. This way, multiple templates can be matched simultaneously at the coarse levels of the search, resulting in various speed-up factors. Furthermore, in matching, features are distinguished by type and separate DTs are computed for each type (e.g. based on edge orientations). These concepts are illustrated in the application of traffic sign detection. Dariu Gavrila |
ICPR | 1 |
| 1996 | 3-D model-based tracking of humans in action: a multi-view approachabstractWe present a vision system for the 3-D model-based tracking of unconstrained human movement. Using image sequences acquired simultaneously from multiple views, we recover the 3-D body pose at each time instant without the use of markers. The pose-recovery problem is formulated as a search problem and entails finding the pose parameters of a graphical human model whose synthesized appearance is most similar to the actual appearance of the real human in the multi-view images. The models used for this purpose are acquired from the images. We use a decomposition approach and a best-first technique to search through the high dimensional pose parameter space. A robust variant of chamfer matching is used as a fast similarity measure between synthesized and real edge images. We present initial tracking results from a large new Humans-in-Action (HIA) database containing more than 2500 frames in each of four orthogonal views. They contain subjects involved in a variety of activities, of various degrees of complexity, ranging from the more simple one-person hand waving to the challenging two-person close interaction in the Argentine Tango. Dariu Gavrila, Larry Davis 0001 |
CVPR | 1 |
| 1996 | Hermite deformable contoursabstractWe propose the Hermite representation for deformable contour finding. This representation compares favorably in terms of versatility and controllability with other local contour representations that have been used previously for this purpose. The Hermite representation allows a compact representation of curved shapes, without the smoothing out of corners. It is also well suited for both interactive and tracking applications. The Hermite representation is used to formulate the contour finding problem as an optimization problem using a maximum a posteriori energy criterion. Optimization is performed by dynamic programming. Our approach to contour tracking decouples the effects of transformation and deformation, using a template matching strategy to robustly account for the transformation effect. We demonstrate these ideas on a variety of images from different domains. Dariu Gavrila |
ICPR | 1 |
| 1992 | 3D object recognition from 2D images using geometric hashing
Dariu Gavrila, Frans C. A. Groen |
Pattern Recognit. Lett. | 1 |