Julian F. P. Kooij

dblp:56/7158 · also Julian Francisco Pieter Kooij · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-9919-0710ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 NeuSEditor: From Multi-View Images to Text-Guided Neural Surface Edits
abstract
Implicit surface representations are valued for their compactness and continuity, but they pose significant challenges for editing. Despite recent advancements, existing methods often fail to preserve identity and maintain geometric consistency during editing. To address these challenges, we present NeuSEditor, a novel method for text-guided editing of neural implicit surfaces derived from multi-view images. NeuSEditor introduces an identity-preserving architecture that efficiently separates scenes into foreground and background, enabling precise modifications without altering the scene-specific elements. Our geometry-aware distillation loss significantly enhances rendering and geometric quality. Our method simplifies the editing workflow by eliminating the need for continuous dataset updates and source prompting. NeuSEditor outperforms recent state-of-the-art methods, delivering superior quantitative and qualitative results. For visual results, visit: neuseditor.github.io.
Nail Ibrahimli, Julian F. P. Kooij, Liangliang Nan
3DV2
2025 VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flow
abstract
Scene flow estimation aims to recover per-point motion from two adjacent LiDAR scans. However, in real-world applications such as autonomous driving, points rarely move independently of others, especially for nearby points belonging to the same object, which often share the same motion. Incorporating this locally rigid motion constraint has been a key challenge in self-supervised scene flow estimation, which is often addressed by post-processing or appending extra regularization. While these approaches are able to improve the rigidity of predicted flows, they lack an architectural inductive bias for local rigidity within the model structure, leading to suboptimal learning efficiency and inferior performance. In contrast, we enforce local rigidity with a lightweight add-on module in neural network design, enabling end-to-end learning. We design a discretized voting space that accommodates all possible translations and then identify the one shared by nearby points by differentiable voting. Additionally, to ensure computational efficiency, we operate on pillars rather than points and learn representative features for voting per pillar. We plug the Voting Module into popular model designs and evaluate its benefit on Argoverse 2 and Waymo datasets. We outperform baseline works with only marginal compute overhead. Code is available at https://github.com/tudelft-iv/VoteFlow.
Yancong Lin, Liangliang Nan, Julian F. P. Kooij, Holger Caesar
CVPR4
2025 A Vehicle System for Navigating Among Vulnerable Road Users Including Remote Operation
abstract
We present a vehicle system capable of navigating safely and efficiently around Vulnerable Road Users (VRUs), such as pedestrians and cyclists. The system comprises key modules for environment perception, localization and mapping, motion planning, and control, integrated into a prototype vehicle. A key innovation is a motion planner based on Topology-driven Model Predictive Control (T-MPC). The guidance layer generates multiple trajectories in parallel, each representing a distinct strategy for obstacle avoidance or non-passing. The underlying trajectory optimization constrains the joint probability of collision with VRUs under generic uncertainties. To address extraordinary situations (“edge cases”) that go beyond the autonomous capabilities — such as construction zones or encounters with emergency responders — the system includes an option for remote human operation, supported by visual and haptic guidance. In simulation, our motion planner outperforms three baseline approaches in terms of safety and efficiency. We also demonstrate the full system in prototype vehicle tests on a closed track, both in autonomous and remotely operated modes.
Oscar de Groot, Alberto Bertipaglia, Hidde J.-H. Boekema, Vishrut Jain, Marcell Kegl, Varun Kotian, Ted de Vries Lentsch, Yancon Lin, Chrysovalanto Messiou, Emma Schippers, Farzam Tajdari, Zimin Xia, Mubariz Zaffar, Ronald M. Ensing, Mario Garzon, Javier Alonso-Mora, Holger Caesar, Laura Ferranti, Riender Happee, Julian F. P. Kooij, Barys N. Shyrokau, Dariu Gavrila
IV21
2024 MuVieCAST: Multi-View Consistent Artistic Style Transfer
abstract
We introduce MuVieCAST, a modular multi-view consistent style transfer network architecture that enables consistent style transfer between multiple viewpoints of the same scene. This network architecture supports both sparse and dense views, making it versatile enough to handle a wide range of multi-view image datasets. The approach consists of three modules that perform specific tasks related to style transfer, namely content preservation, image transformation, and multi-view consistency enforcement. We extensively evaluate our approach across multiple application domains including depth-map-based point cloud fusion, mesh reconstruction, and novel-view synthesis. Our experiments reveal that the proposed framework achieves an exceptional generation of stylized images, exhibiting consistent outcomes across perspectives. A user study focusing on novel-view synthesis further confirms these results, with approximately $68 \%$ of cases participants expressing a preference for our generated outputs compared to the recent state-of-the-art method. Our modular framework is extensible and can easily be integrated with various backbone architectures, making it a flexible solution for multi-view style transfer. More results are demonstrated on our project page: muviecast.github.io
Nail Ibrahimli, Julian F. P. Kooij, Liangliang Nan
3DV2
2024 On the Estimation of Image-Matching Uncertainty in Visual Place Recognition
abstract
In Visual Place Recognition (VPR) the pose of a query image is estimated by comparing the image to a map of reference images with known reference poses. As is typical for image retrieval problems, a feature extractor maps the query and reference images to a feature space, where a nearest neighbor search is then performed. However, till recently little attention has been given to quantifying the confidence that a retrieved reference image is a correct match. Highly certain but incorrect retrieval can lead to catastrophic failure of VPR-based localization pipelines. This work compares for the first time the main approaches for estimating the image-matching uncertainty, including the traditional retrieval-based uncertainty estimation, more recent data-driven aleatoric uncertainty estimation, and the compute-intensive geometric verification. We further formulate a simple baseline method, “SUE”, which unlike the other methods considers the freely-available poses of the reference images in the map. Our experiments reveal that a simple L2-distance between the query and reference descriptors is already a better estimate of image-matching uncertainty than current data-driven approaches. SUE outperforms the other efficient uncertainty estimation methods, and its uncertainty estimates complement the computationally expensive geometric verification approach. Future works for uncertainty estimation in VPR should consider the baselines discussed in this work.
Mubariz Zaffar, Liangliang Nan, Julian F. P. Kooij
CVPR3
2024 Adapting Fine-Grained Cross-View Localization to Areas Without Fine Ground Truth
Zimin Xia, Yujiao Shi 0002, Hongdong Li, Julian F. P. Kooij
ECCV (31)4
2024 UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
abstract
Multi-sensor object detection is an active research topic in automated driving, but the robustness of such detection models against missing sensor input (modality missing), e.g., due to a sudden sensor failure, is a critical problem which remains under-studied. In this work, we propose UniBEV, an end-to-end multi-modal 3D object detection framework designed for robustness against missing modalities: UniBEV can operate on LiDAR plus camera input, but also on LiDAR-only or camera-only input without retraining. To facilitate its detector head to handle different input combinations, UniBEV aims to create well-aligned Bird’s Eye View (BEV) feature maps from each available modality. Unlike prior BEV-based multi-modal detection methods, all sensor modalities follow a uniform approach to resample features from the original sensor coordinate systems to the BEV features. We furthermore investigate the robustness of various fusion strategies w.r.t. missing modalities: the commonly used feature concatenation, but also channel-wise averaging, and a generalization to weighted averaging termed Channel Normalized Weights. To validate its effectiveness, we compare UniBEV to state-of-the-art BEVFusion and MetaBEV on nuScenes over all sensor input combinations. In this setting, UniBEV achieves better performance than these baselines for all input combinations. An ablation study shows the robustness benefits of fusing by weighted averaging over regular concatenation, and of sharing queries between the BEV encoders of each modality. Our code is available at https://github.com/tudelft-iv/UniBEV.
Holger Caesar, Liangliang Nan, Julian F. P. Kooij
IV4
2024 Convolutional Cross-View Pose Estimation
abstract
We propose a novel end-to-end method for cross-view pose estimation. Given a ground-level query image and an aerial image that covers the query's local neighborhood, the 3 Degrees-of-Freedom camera pose of the query is estimated by matching its image descriptor to descriptors of local regions within the aerial image. The orientation-aware descriptors are obtained by using a translationally equivariant convolutional ground image encoder and contrastive learning. The Localization Decoder produces a dense probability distribution in a coarse-to-fine manner with a novel Localization Matching Upsampling module. A smaller Orientation Decoder produces a vector field to condition the orientation estimate on the localization. Our method is validated on the VIGOR and KITTI datasets, where it surpasses the state-of-the-art baseline by 72% and 36% in median localization error for comparable orientation estimation accuracy. The predicted probability distribution can represent localization ambiguity, and enables rejecting possible erroneous predictions. Without re-training, the model can infer on ground images with different field of views and utilize orientation priors if available. On the Oxford RobotCar dataset, our method can reliably estimate the ego-vehicle's pose over time, achieving a median localization error under 1 m and a median orientation error of around 1$^{\circ }$at 14 FPS.
Zimin Xia, Olaf Booij, Julian F. P. Kooij
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation
abstract
This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists of ground and aerial feature extractors, feature aggregators, and a pose predictor. The feature extractors extract dense features from the ground and aerial images. Given a set of candidate camera poses, the feature aggregators construct a single ground descriptor and a set of pose-dependent aerial descriptors. Notably, our novel aerial feature aggregator has a cross-view attention module for ground-view guided aerial feature selection and utilizes the geometric projection of the ground camera's viewing frustum on the aerial image to pool features. The efficient construction of aerial descriptors is achieved using precomputed masks. SliceMatch is trained using contrastive learning and pose estimation is formulated as a similarity comparison between the ground descriptor and the aerial descriptors. Compared to the state-of-the-art, SliceMatch achieves a 19% lower median localization error on the VIGOR benchmark using the same VGG16 backbone at 150 frames per second, and a 50% lower error when using a ResNet50 backbone.
Ted de Vries Lentsch, Zimin Xia, Holger Caesar, Julian F. P. Kooij
CVPR4
2023 How Informative is the Approximation Error from Tensor Decomposition for Neural Network Compression?
Jetze Schuurmans, Kim Batselier, Julian F. P. Kooij
ICLR3
2023 CoPR: Toward Accurate Visual Localization With Continuous Place-Descriptor Regression
abstract
Visual place recognition (VPR) is an image-based localization method that estimates the camera location of a query image by retrieving the most similar reference image from a map of geo-tagged reference images. In this work, we look into two fundamental bottlenecks for its localization accuracy: 1) reference map sparseness and 2) viewpoint invariance. First, the reference images for VPR are only available at sparse poses in a map, which enforces an upper bound on the maximum achievable localization accuracy through VPR. We, therefore, propose Continuous Place-descriptor Regression (CoPR) to densify the map and improve localization accuracy. We study various interpolation and extrapolation models to regress additional VPR feature descriptors from only the existing references. Second, we compare different feature encoders and show that CoPR presents value for all of them. We evaluate our models on three existing public datasets and report on average around 30% improvement in VPR-based localization accuracy using CoPR, on top of the 15% increase by using a viewpoint-variant loss for the feature encoder. The complementary relation between CoPR and relative pose estimation is also discussed.
Mubariz Zaffar, Liangliang Nan, Julian F. P. Kooij
IEEE Trans. Robotics3
2022 Push-the-Boundary: Boundary-aware Feature Propagation for Semantic Segmentation of 3D Point Clouds
abstract
Feedforward fully convolutional neural networks currently dominate in semantic segmentation of 3D point clouds. Despite their great success, they suffer from the loss of local information at low-level layers, posing significant challenges to accurate scene segmentation and precise object boundary delineation. Prior works either address this issue by post-processing or jointly learn object boundaries to implicitly improve feature encoding of the networks. These approaches often require additional modules which are difficult to integrate into the original architecture. To improve the segmentation near object boundaries, we propose a boundary-aware feature propagation mechanism. This mechanism is achieved by exploiting a multitask learning framework that aims to explicitly guide the boundaries to their original locations. With one shared encoder, our network outputs (i) boundary localization, (ii) prediction of directions pointing to the object's interior, and (iii) semantic segmentation, in three parallel streams. The predicted boundaries and directions are fused to propagate the learned features to refine the segmentation. We conduct extensive experiments on the S3DIS and SensatUrban datasets against various baseline methods, demonstrating that our proposed approach yields consistent improvements by reducing boundary errors. Our code is available at https://github.com/shenglandu/PushBoundary.
Shenglan Du, Nail Ibrahimli, Jantien E. Stoter, Julian F. P. Kooij, Liangliang Nan
3DV4
2022 Visual Cross-View Metric Localization with Dense Uncertainty Estimates
Zimin Xia, Olaf Booij, Marco Manfredi, Julian F. P. Kooij
ECCV (39)4
2022 BackboneAnalysis: Structured Insights into Compute Platforms from CNN Inference Latency
abstract
Customization of a convolutional neural network (CNN) to a specific compute platform involves finding an optimal pareto state between computational complexity of the CNN and resulting throughput in operations per second on the compute platform. However, existing inference performance benchmarks compare complete backbones that entail many differences between their CNN configurations, which do not provide insights in how fine-grade layer design choices affect this balance.We present BackboneAnalysis, a methodology for extracting structured insights into the trade-off for a chosen target compute platform. Within a one-factor-at-a-time analysis setup, CNN architectures are systematically varied and evaluated based on throughput and latency measurements irrespective of model accuracy. Thereby, we investigate the configuration factors input shape, batch size, kernel size and convolutional layer type.In our experiments, we deploy BackboneAnalysis on a Xavier iGPU and a Coral Edge TPU accelerator. The analysis reveals that the general assumption from optimal Roofline performance that higher operation density in CNNs leads to higher throughput does not always hold. These results highlight the importance for a neural network architect to be aware of platform-specific latency and throughput behavior in order to derive sensible configuration decisions for a custom CNN.
Frank Hafner, Matthias Zeller, Mark Schutera, Jochen Abhau, Julian F. P. Kooij
IV5
2022 Cross-modal distillation for RGB-depth person re-identification
Frank Hafner, Amran Bhuyian, Julian F. P. Kooij, Eric Granger
Comput. Vis. Image Underst.3
2021 VPR-Bench: An Open-Source Visual Place Recognition Evaluation Framework with Quantifiable Viewpoint and Appearance Change
abstract
Abstract Visual place recognition (VPR) is the process of recognising a previously visited place using visual information, often under varying appearance conditions and viewpoint changes and with computational constraints. VPR is related to the concepts of localisation, loop closure, image retrieval and is a critical component of many autonomous navigation systems ranging from autonomous vehicles to drones and computer vision systems. While the concept of place recognition has been around for many years, VPR research has grown rapidly as a field over the past decade due to improving camera hardware and its potential for deep learning-based techniques, and has become a widely studied topic in both the computer vision and robotics communities. This growth however has led to fragmentation and a lack of standardisation in the field, especially concerning performance evaluation. Moreover, the notion of viewpoint and illumination invariance of VPR techniques has largely been assessed qualitatively and hence ambiguously in the past. In this paper, we address these gaps through a new comprehensive open-source framework for assessing the performance of VPR techniques, dubbed “VPR-Bench”. VPR-Bench (Open-sourced at: https://github.com/MubarizZaffar/VPR-Bench ) introduces two much-needed capabilities for VPR researchers: firstly, it contains a benchmark of 12 fully-integrated datasets and 10 VPR techniques, and secondly, it integrates a comprehensive variation-quantified dataset for quantifying viewpoint and illumination invariance. We apply and analyse popular evaluation metrics for VPR from both the computer vision and robotics communities, and discuss how these different metrics complement and/or replace each other, depending upon the underlying applications and system requirements. Our analysis reveals that no universal SOTA VPR technique exists, since: (a) state-of-the-art (SOTA) performance is achieved by 8 out of the 10 techniques on at least one dataset, (b) SOTA technique in one community does not necessarily yield SOTA performance in the other given the differences in datasets and metrics. Furthermore, we identify key open challenges since: (c) all 10 techniques suffer greatly in perceptually-aliased and less-structured environments, (d) all techniques suffer from viewpoint variance where lateral change has less effect than 3D change, and (e) directional illumination change has more adverse effects on matching confidence than uniform illumination change. We also present detailed meta-analyses regarding the roles of varying ground-truths, platforms, application requirements and technique parameters. Finally, VPR-Bench provides a unified implementation to deploy these VPR techniques, metrics and datasets, and is extensible through templates.
Mubariz Zaffar, Sourav Garg, Michael Milford, Julian F. P. Kooij, David Flynn, Klaus D. McDonald-Maier, Shoaib Ehsan
Int. J. Comput. Vis.4
2020 End-to-End Learning of Decision Trees and Forests
abstract
Abstract Conventional decision trees have a number of favorable properties, including a small computational footprint, interpretability, and the ability to learn from little training data. However, they lack a key quality that has helped fuel the deep learning revolution: that of being end-to-end trainable. Kontschieder et al. (ICCV, 2015) have addressed this deficit, but at the cost of losing a main attractive trait of decision trees: the fact that each sample is routed along a small subset of tree nodes only. We here present an end-to-end learning scheme for deterministic decision trees and decision forests. Thanks to a new model and expectation–maximization training scheme, the trees are fully probabilistic at train time, but after an annealing process become deterministic at test time. In experiments we explore the effect of annealing visually and quantitatively, and find that our method performs on par or superior to standard learning algorithms for oblique decision trees and forests. We further demonstrate on image datasets that our approach can learn more complex split functions than common oblique ones, and facilitates interpretability through spatial regularization.
Thomas M. Hehn, Julian F. P. Kooij, Fred A. Hamprecht
Int. J. Comput. Vis.2
2019 RGB-Depth Cross-Modal Person Re-identification
abstract
Person re-identification is a key challenge for surveillance across multiple sensors. Prompted by the advent of powerful deep learning models for visual recognition, and inexpensive RGBD cameras and sensor-rich mobile robotic platforms, e.g. self-driving vehicles, we investigate the relatively unexplored problem of cross-modal re-identification of persons between RGB (color) and depth images. The considerable divergence in data distributions across different sensor modalities introduces additional challenges to the typical difficulties like distinct viewpoints, occlusions, and pose and illumination variation. While some work has investigated re-identification across RGB and infrared, we take inspiration from successes in transfer learning from RGB to depth in object detection tasks. Our main contribution is a novel cross-modal distillation network for robust person re-identification, which learns a shared feature representation space of person's appearance in both RGB and depth images. The proposed network was compared to conventional and deep learning approaches proposed for other cross-domain re-identification tasks. Results obtained on the public BIWI and RobotPKU datasets indicate that the proposed method can significantly outperform the state-of-the-art approaches by up to 10.5% mAp, demonstrating the benefit of the proposed distillation paradigm.
Frank Hafner, Amran Bhuiyan, Julian F. P. Kooij, Eric Granger
AVSS3
2019 An Extrinsic Calibration Tool for Radar, Camera and Lidar
abstract
We present a novel open-source tool for extrinsic calibration of radar, camera and lidar. Unlike currently available offerings, our tool facilitates joint extrinsic calibration of all three sensing modalities on multiple measurements. Furthermore, our calibration target design extends existing work to obtain simultaneous measurements for all these modalities. We study how various factors of the calibration procedure affect the outcome on real multi-modal measurements of the target. Three different configurations of the optimization criterion are considered, namely using error terms for a minimal amount of sensor pairs, or using terms for all sensor pairs with additional loop closure constraints, or by adding terms for structure estimation in a probabilistic model. The experiments further evaluate how the number of calibration boards affect calibration performance, and robustness against different levels of zero mean Gaussian noise. Our results show that all configurations achieve good results for lidar to camera errors and that fully connected pose estimation shows the best performance for lidar to radar errors when more than five board locations are used.
Joris Domhof, Julian F. P. Kooij, Dariu M. Kooij
ICRA2
2019 SafeVRU: A Research Platform for the Interaction of Self-Driving Vehicles with Vulnerable Road Users
abstract
This paper presents our research platform Safe VRU for the interaction of self-driving vehicles with Vulnerable Road Users (VRUs, i.e., pedestrians and cyclists). The paper details the design (implemented with a modular structure within ROS) of the full stack of vehicle localization, environment perception, motion planning, and control, with emphasis on the environment perception and planning modules. The environment perception detects the VRUs using a stereo camera and predicts their paths with Dynamic Bayesian Networks (DBNs), which can account for switching dynamics. The motion planner is based on model predictive contouring control (MPCC) and takes into account vehicle dynamics, control objectives (e.g., desired speed), and perceived environment (i.e., the predicted VRU paths with behavioral uncertainties) over a certain time horizon. We present simulation and real-world results to illustrate the ability of our vehicle to plan and execute collision-free trajectories in the presence of VRUs.
Laura Ferranti, Bruno Brito, Ewoud A. I. Pool, Ronald M. Ensing, Riender Happee, Barys N. Shyrokau, Julian F. P. Kooij, Javier Alonso-Mora, Dariu Gavrila
IV8
2019 Instance Stixels: Segmenting and Grouping Stixels into Objects
abstract
State-of-the-art stixel methods fuse dense stereo and semantic class information, e.g. from a Convolutional Neural Network (CNN), into a compact representation of driveable space, obstacles, and background. However, they do not explicitly differentiate instances within the same class. We investigate several ways to augment single-frame stixels with instance information, which can similarly be extracted by a CNN from the color input. As a result, our novel Instance Stixels method efficiently computes stixels that do account for boundaries of individual objects, and represents individual instances as grouped stixels that express connectivity. Experiments on Cityscapes demonstrate that including instance information into the stixel computation itself, rather than as a post-processing step, increases Instance AP performance with approximately the same number of stixels. Qualitative results confirm that segmentation improves, especially for overlapping objects of the same class. Additional tests with ground truth instead of CNN output show that the approach has potential for even larger gains. Our Instance Stixels software is made freely available for non-commercial research purposes.
Thomas M. Hehn, Julian F. P. Kooij, Dariu Gavrila
IV2
2019 Occlusion aware sensor fusion for early crossing pedestrian detection
abstract
Early and accurate detection of crossing pedestrians is crucial in automated driving to execute emergency manoeuvres in time. This is a challenging task in urban scenarios however, where people are often occluded (not visible) behind objects, e.g. other parked vehicles. In this paper, an occlusion aware multi-modal sensor fusion system is proposed to address scenarios with crossing pedestrians behind parked vehicles. Our proposed method adjusts the detection rate in different areas based on sensor visibility. We argue that using this occlusion information can help to evaluate the measurements. Our experiments on real world data show that fusing radar and stereo camera for such tasks is beneficial, and that including occlusion into the model helps to detect pedestrians earlier and more accurately.
Andras Palffy, Julian F. P. Kooij, Dariu Gavrila
IV2
2019 Context-based cyclist path prediction using Recurrent Neural Networks
abstract
This paper proposes a Recurrent Neural Network (RNN) for cyclist path prediction to learn the effect of contextual cues on the behavior directly in an end- to-end approach, removing the need for any annotations. The proposed RNN incorporates three distinct contextual cues: one related to actions of the cyclist, one related to the location of the cyclist on the road, and one related to the interaction between the cyclist and the egovehicle. The RNN predicts a Gaussian distribution over the future position of the cyclist one second into the future with a higher accuracy, compared to a current state-of-the-art model that is based on dynamic mode annotations, where our model attains an average prediction error of 33 cm one second into the future.
Ewoud A. I. Pool, Julian F. P. Kooij, Dariu Gavrila
IV2
2019 Context-Based Path Prediction for Targets with Switching Dynamics
abstract
Anticipating future situations from streaming sensor data is a key perception challenge for mobile robotics and automated vehicles. We address the problem of predicting the path of objects with multiple dynamic modes. The dynamics of such targets can be described by a Switching Linear Dynamical System (SLDS). However, predictions from this probabilistic model cannot anticipate when a change in dynamic mode will occur. We propose to extract various types of cues with computer vision to provide context on the target’s behavior, and incorporate these in a Dynamic Bayesian Network (DBN). The DBN extends the SLDS by conditioning the mode transition probabilities on additional context states. We describe efficient online inference in this DBN for probabilistic path prediction, accounting for uncertainty in both measurements and target behavior. Our approach is illustrated on two scenarios in the Intelligent Vehicles domain concerning pedestrians and cyclists, so-called Vulnerable Road Users (VRUs). Here, context cues include the static environment of the VRU, its dynamic environment, and its observed actions. Experiments using stereo vision data from a moving vehicle demonstrate that the proposed approach results in more accurate path prediction than SLDS at the relevant short time horizon (1 s). It slightly outperforms a computationally more demanding state-of-the-art method.
Julian F. P. Kooij, Fabian Flohr, Ewoud A. I. Pool, Dariu Gavrila
Int. J. Comput. Vis.1
2017 Using road topology to improve cyclist path prediction
abstract
We learn motion models for cyclist path prediction on real-world tracks obtained from a moving vehicle, and propose to exploit the local road topology to obtain better predictive distributions. The tracks are extracted from the Tsinghua-Daimler Cyclist Benchmark for cyclist detection, and corrected for vehicle egomotion. Tracks are then spatially aligned to local curves and crossings in the road. We study a standard approach for path prediction in the literature based on Kalman Filters, as well as a mixture of specialized filters related to specific road orientations at junctions. Our experiments demonstrate an improved prediction accuracy (up to 20% on sharp turns) of mixing specialized motion models for canonical directions, and prior knowledge on the road topology. The new track data complements the existing video, disparity and annotation data of the original benchmark, and will be made publicly available.
Ewoud A. I. Pool, Julian F. P. Kooij, Dariu Gavrila
Intelligent Vehicles Symposium2
2016 Depth-Aware Motion Magnification
Julian F. P. Kooij, Jan C. van Gemert
ECCV (8)1
2016 SenseCap: Synchronized Data Collection with Microsoft Kinect2 and LeapMotion
abstract
We present a new recording tool to capture synchronized video and skeletal data streams from cheap sensors such as the Microsoft Kinect2, and LeapMotion. While other recording tools act as virtual playback devices for testing on-line real-time applications, we target multi-media data collection for off-line processing. Images are encoded in common video formats, and skeletal data as flat text tables. This approach enables long duration recordings (e.g. over 30 minutes), and supports post-hoc mapping of the Kinect2 depth video to the color space if needed. By using common file formats, the data can be played back and analyzed on any other computer, without requiring sensor specific SDKs to be installed. The project is released under a 3-clause BSD license, and consists of an extensible C++11 framework, with support for the official Microsoft Kinect 2 and LeapMotion APIs to record, a command-line interface, and a Matlab GUI to initiate, inspect, and load Kinect2 recordings.
Julian F. P. Kooij
ACM Multimedia1
2016 Multi-modal human aggression detection
abstract
This paper presents a smart surveillance system named CASSANDRA, aimed at detecting instances of aggressive human behavior in public environments. A distinguishing aspect of CASSANDRA is the exploitation of complementary audio and video cues to disambiguate scene activity in real-life environments. From the video side, the system uses overlapping cameras to track persons in 3D and to extract features regarding the limb motion relative to the torso. From the audio side, it classifies instances of speech, screaming, singing, and kicking-object. The audio and video cues are fused with contextual cues (interaction, auxiliary objects); a Dynamic Bayesian Network (DBN) produces an estimate of the ambient aggression level. Our prototype system is validated on a realistic set of scenarios performed by professional actors at an actual train station to ensure a realistic audio and video noise setting.
Julian F. P. Kooij, Martijn Liem, Johannes D. Krijnders, Tjeerd C. Andringa, Dariu Gavrila
Comput. Vis. Image Underst.1
2016 Mixture of Switching Linear Dynamics to Discover Behavior Patterns in Object Tracks
abstract
We present a novel non-parametric Bayesian model to jointly discover the dynamics of low-level actions and high-level behaviors of tracked objects. In our approach, actions capture both linear, low-level object dynamics, and an additional spatial distribution on where the dynamic occurs. Furthermore, behavior classes capture high-level temporal motion dependencies in Markov chains of actions, thus each learned behavior is a switching linear dynamical system. The number of actions and behaviors is discovered from the data itself using Dirichlet Processes. We are especially interested in cases where tracks can exhibit large kinematic and spatial variations, e.g. person tracks in open environments, as found in the visual surveillance and intelligent vehicle domains. The model handles real-valued features directly, so no information is lost by quantizing measurements into 'visual words', and variations in standing, walking and running can be discovered without discrete thresholds. We describe inference using Markov Chain Monte Carlo sampling and validate our approach on several artificial and real-world pedestrian track datasets from the surveillance and intelligent vehicle domain. We show that our model can distinguish between relevant behavior patterns that an existing state-of-the-art hierarchical model for clustering and simpler model variants cannot. The software and the artificial and surveillance datasets are made publicly available for benchmarking purposes.
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Identifying multiple objects from their appearance in inaccurate detections
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila
Comput. Vis. Image Underst.1
2015 A Probabilistic Framework for Joint Pedestrian Head and Body Orientation Estimation
abstract
We present a probabilistic framework for the joint estimation of pedestrian head and body orientation from a mobile stereo vision platform. For both head and body parts, we convert the responses of a set of orientation-specific detectors into a (continuous) probability density function. The parts are localized by means of apictorial structureapproach, which balances part-based detector responses with spatial constraints. Head and body orientations are estimated jointly to account for anatomical constraints. The joint single-frame orientation estimates are integrated over time by particle filtering. The experiments involved data from a vehicle-mounted stereo vision camera in a realistic traffic setting; 65 pedestrian tracks were supplied by a state-of-the-art pedestrian tracker. We show that the proposed joint probabilistic orientation estimation framework reduces the mean absolute head and body orientation error up to 15° compared with simpler methods. This results in a mean absolute head/body orientation error of about 21°/19°, which remains fairly constant up to a distance of 25 m. Our system currently runs in near real time (8–9 Hz).
Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila
IEEE Trans. Intell. Transp. Syst.3
2014 Context-Based Pedestrian Path Prediction
Julian F. P. Kooij, Nicolas Schneider, Fabian Flohr, Dariu Gavrila
ECCV (6)1
2014 Joint probabilistic pedestrian head and body orientation estimation
abstract
We present an approach for the joint probabilistic estimation of pedestrian head and body orientation in the context of intelligent vehicles. For both, head and body, we convert the output of a set of orientation-specific detectors into a full (continuous) probability density function. The parts are localized with a pictorial structure approach which balances part-based detector output with spatial constraints. Head and body orientation estimates are furthermore coupled probabilistically to account for anatomical constraints. Finally, the coupled single-frame orientation estimates are integrated over time by particle filtering. The experiments involve 37 pedestrian tracks obtained from an external stereo vision-based pedestrian detector in realistic traffic settings. We show that the proposed joint probabilistic orientation estimation approach reduces the mean head and body orientation error by 10 degrees and more.
Fabian Flohr, Madalin Dumitru-Guzu, Julian F. P. Kooij, Dariu Gavrila
Intelligent Vehicles Symposium3
2014 Analysis of pedestrian dynamics from a vehicle perspective
abstract
Accurate motion models are key to many tasks in the intelligent vehicle domain, but simple Linear Dynamics (e.g. Kalman filtering) do not exploit the spatio-temporal context of motion. We present a method to learn Switching Linear Dynamics of object tracks observed from within a driving vehicle. Each switching state captures object dynamics as a mean motion with variance, but also has an additional spatial distribution on where the dynamic is seen relative to the vehicle. Thus, both an object's previous movements and current location will make certain dynamics more probable for subsequent time steps. To train the model, we use Bayesian inference to sample parameters from the posterior, and jointly learn the required number of dynamics. Unlike Maximum Likelihood learning, inference is robust against overfitting and poor initialization. We demonstrate our approach on an ego-motion compensated track dataset of pedestrians, and illustrate how the switching dynamics can make more accurate path predictions than a mixture of linear dynamics for crossing pedestrians.
Julian F. P. Kooij, Nicolas Schneider, Dariu Gavrila
Intelligent Vehicles Symposium1
2012 A Non-parametric Hierarchical Model to Discover Behavior Dynamics from Tracks
Julian F. P. Kooij, Gwenn Englebienne, Dariu Gavrila
ECCV (6)1