Gijs Dubbelman

dblp:45/2051 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0001-6635-3245ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Systems, architecture and hardware · 7 · 7 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Your ViT is Secretly an Image Segmentation Model
abstract
Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multi-scale features, a pixel decoder to fuse these features, and a Transformer decoder that uses the fused features to make predictions. In this paper, we show that the inductive biases introduced by these task-specific components can instead be learned by the ViT itself, given sufficiently large models and extensive pre-training. Based on these findings, we introduce the Encoder-Only Mask Transformer (EoMT), which repurposes the plain ViT architecture to conduct image segmentation. With large-scale models and pre-training, EoMT obtains a segmentation accuracy similar to state-of-the-art models that use task-specific components. At the same time, EoMT is significantly faster than these methods due to its architectural simplicity, e.g., up to 4 × faster with ViT-L. Across a range of model sizes, EoMT demonstrates an optimal balance between segmentation accuracy and prediction speed, suggesting that compute resources are better spent on scaling the ViT itself rather than adding architectural complexity. Code: https://www.tue-mps.org/eomt/.
Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans, Narges Norouzi, Giuseppe Averta, Bastian Leibe, Gijs Dubbelman, Daan de Geus
CVPR7
2024 Task-Aligned Part-Aware Panoptic Segmentation Through Joint Object-Part Representations
abstract
Part-aware panoptic segmentation (PPS) requires (a) that each foreground object and background region in an image is segmented and classified, and (b) that all parts within foreground objects are segmented, classified and linked to their parent object. Existing methods approach PPS by separately conducting object-level and part-level segmentation. However, their part-level predictions are not linked to individual parent objects. Therefore, their learning objective is not aligned with the PPS task objective, which harms the PPS performance. To solve this, and make more accurate PPS predictions, we propose Task-Aligned Part-aware Panoptic Segmentation (TAPPS). This method uses a set of shared queries to jointly predict (a) object-level segments, and (b) the part-level segments within those same objects. As a result, TAPPS learns to predict part-level segments that are linked to individual parent objects, aligning the learning objective with the task objective, and allowing TAPPS to leverage joint object-part representations. With experiments, we show that TAPPS considerably outperforms methods that predict objects and parts separately, and achieves new state-of-the-art PPS results.
Daan de Geus, Gijs Dubbelman
CVPR2
2024 ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
abstract
This work presents Adaptive Local-then-Global Merging (ALGM), a token reduction method for semantic segmentation networks that use plain Vision Transformers. ALGM merges tokens in two stages: (1) In the first network layer, it merges similar tokens within a small local window and (2) halfway through the network, it merges similar tokens across the entire image. This is motivated by an analysis in which we found that, in those situations, tokens with a high cosine similarity can likely be merged without a drop in segmentation quality. With extensive experiments across multiple datasets and network configurations, we show that ALGM not only significantly improves the throughput by up to 100%, but can also enhance the mean IoU by up to + 1.1, thereby achieving a better trade-off between segmentation quality and efficiency than existing methods. Moreover, our approach is adaptive during inference, meaning that the same model can be used for optimal efficiency or accuracy, depending on the application. Code is available at https://tue-mps.github.io/ALGM.
Narges Norouzi, Svetlana Orlova, Daan de Geus, Gijs Dubbelman
CVPR4
2024 Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning
abstract
Learning anticipation in Multi-Agent Reinforcement Learning (MARL) is a reasoning paradigm where agents anticipate the learning steps of other agents to improve cooperation among themselves. As MARL uses gradient-based optimization, learning anticipation requires using Higher-Order Gradients (HOG), with so-called HOG methods. Existing HOG methods are based on policy parameter anticipation, i.e., agents anticipate the changes in policy parameters of other agents. Currently, however, these existing HOG methods have only been developed for differentiable games or games with small state spaces. In this work, we demonstrate that in the case of non-differentiable games with large state spaces, existing HOG methods do not perform well and are inefficient due to their inherent limitations related to policy parameter anticipation and multiple sampling stages. To overcome these problems, we propose Off-Policy Action Anticipation (OffPA2), a novel framework that approaches learning anticipation through action anticipation, i.e., agents anticipate the changes in actions of other agents, via off-policy sampling. We theoretically analyze our proposed OffPA2 and employ it to develop multiple HOG methods that are applicable to non-differentiable games with large state spaces. We conduct a large set of experiments and illustrate that our proposed HOG methods outperform the existing ones regarding efficiency and performance.
Ariyan Bighashdel, Daan de Geus, Pavol Jancura, Gijs Dubbelman
J. Mach. Learn. Res.4
2023 Content-aware Token Sharing for Efficient Semantic Segmentation with Vision Transformers
abstract
This paper introduces Content-aware Token Sharing (CTS), a token reduction approach that improves the computational efficiency of semantic segmentation networks that use Vision Transformers (ViTs). Existing works have proposed token reduction approaches to improve the efficiency of ViT-based image classification networks, but these methods are not directly applicable to semantic segmentation, which we address in this work. We observe that, for semantic segmentation, multiple image patches can share a token if they contain the same semantic class, as they contain redundant information. Our approach leverages this by employing an efficient, class-agnostic policy network that predicts if image patches contain the same semantic class, and lets them share a token if they do. With experiments, we explore the critical design choices of CTS and show its effectiveness on the ADE20K, Pascal Context and Cityscapes datasets, various ViT backbones, and different segmentation decoders. With Content-aware Token Sharing, we are able to reduce the number of processed tokens by up to 44%, without diminishing the segmentation quality.
Chenyang Lu 0002, Daan de Geus, Gijs Dubbelman
CVPR3
2023 Unified Pedestrian Path Prediction Framework: A Comparison Study
abstract
Pedestrian path prediction is an emerging and crucial task in numerous applications, such as autonomous vehicles. Due to the complexity of the task, various formulations are proposed throughout the literature. However, the interconnection between these formulations remains to be seen, which makes a fair comparison challenging. This work proposes a unified pedestrian path prediction framework via Markov decision process (MDP). We demonstrate that by carefully designing the components of the MDP, various standard formulations can be perceived as specific combinations of settings in our framework. Additionally, the unified framework allows us to discover new combinations of settings that integrate the benefits of current formulations improving the prediction performance. We conduct a comparison study and evaluate several formulations in well-controlled experiments. Furthermore, we carefully assess the influence of various settings, such as policy stochasticity and sequential decision-making, on prediction performance. The goal of this work is not to propose a new state-of-the- art method but to study various formulations of the pedestrian path prediction task under a unifying framework and uncover new directions that can eventually advance the current state-of-the-art.
Jarl L. A. Lemmens, Ariyan Bighashdel, Pavol Jancura, Gijs Dubbelman
IV4
2023 Intra-Batch Supervision for Panoptic Segmentation on High-Resolution Images
abstract
Unified panoptic segmentation methods are achieving state-of-the-art results on several datasets. To achieve these results on high-resolution datasets, these methods apply crop-based training. In this work, we find that, although crop-based training is advantageous in general, it also has a harmful side-effect. Specifically, it limits the ability of unified networks to discriminate between large object instances, causing them to make predictions that are confused between multiple instances. To solve this, we propose Intra-Batch Supervision (IBS), which improves a network’s ability to discriminate between instances by introducing additional supervision using multiple images from the same batch. We show that, with our IBS, we successfully address the confusion problem and consistently improve the performance of unified networks. For the high-resolution Cityscapes and Mapillary Vistas datasets, we achieve improvements of up to +2.5 on the Panoptic Quality for thing classes, and even more considerable gains of up to +5.8 on both the pixel accuracy and pixel precision, which we identify as better metrics to capture the confusion problem.
Daan de Geus, Gijs Dubbelman
WACV2
2023 Empirical Generalization Study: Unsupervised Domain Adaptation vs. Domain Generalization Methods for Semantic Segmentation in the Wild
abstract
For autonomous vehicles and mobile robots to safely operate in the real world, i.e., the wild, scene understanding models should perform well in the many different scenarios that can be encountered. In reality, these scenarios are not all represented in the model’s training data, leading to poor performance. To tackle this, current training strategies attempt to either exploit additional unlabeled data with unsupervised domain adaptation (UDA), or to reduce overfitting using the limited available labeled data with domain generalization (DG). However, it is not clear from current literature which of these methods allows for better generalization to unseen data from the wild. Therefore, in this work, we present an evaluation framework in which the generalization capabilities of state-of-the-art UDA and DG methods can be compared fairly. From this evaluation, we find that UDA methods, which leverage unlabeled data, outperform DG methods in terms of generalization, and can deliver similar performance on unseen data as fully-supervised training methods that require all data to be labeled. We show that semantic segmentation performance can be increased up to 30% for a priori unknown data without using any extra labeled data.
Fabrizio J. Piva, Daan de Geus, Gijs Dubbelman
WACV3
2023 Exploiting image translations via ensemble self-supervised learning for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to improve the generalization capacity of models when they are tested on a real-world target domain by learning a model on a source labeled domain. Recently, a UDA method was proposed that addresses the adaptation problem by combining ensemble learning with self-supervised learning. However, this method uses only the source domain to pretrain the model and employs a limited amount of classifiers to create target pseudo labels. To mitigate these deficiencies, in this work, we explore the usage of image translations in combination with ensemble learning and self-supervised learning. To increase the model’s exposure to more variable pretraining data, our method creates multiple diverse image translations, which encourages the learning of domain-invariant features, desired to increase generalization. With these image translations, we are able to learn translation-specific classifiers, which also allows to maximize the amount of ensemble’s classifiers resulting in more robust target pseudo labels. In addition, we propose to use the target domain in pretraining stage to mitigate source domain bias in the network. We evaluate our method on the standard UDA benchmarks, i.e., adapting GTA V and Synthia to Cityscapes, and achieve state-of-the-art results on the mIoU metric. Extensive ablation experiments are reported to highlight the advantageous properties of our UDA strategy.
Fabrizio J. Piva, Gijs Dubbelman
Comput. Vis. Image Underst.2
2023 Correction to: Model-free inverse reinforcement learning with multi-intention, unlabeled, and overlapping demonstrations
abstract
Correction to: Machine Learning https://doi.org/10.1007/s10994-022-06273-x There are two mistakes in the published article: 1. One of the references in the manuscript is incorrect. Here is the incorrect reference: “Bighashdel, A., Meletis, P., Jancura, P., & Dubbelman, G. (2021). In Proceeding of Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 206–221).” and here is the corrected reference: “Bighashdel, A., Meletis, P., Jancura, P., & Dubbelman, G. (2021). Deep Adaptive Multi-Intention Inverse Reinforcement Learning. In Proceedings of Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 206– 221).” 2. There is a missing character “i” in equation 3 of the published manuscript. Here is the incorrect equation:(Formula Presented.) and here is the corrected equation: (Formula Presented) The character “i” is missing in the term “ (a|s) ”. The corrected term is “ (a|s, i)”. The original article has been corrected.
Ariyan Bighashdel, Pavol Jancura, Gijs Dubbelman
Mach. Learn.3
2023 Model-free inverse reinforcement learning with multi-intention, unlabeled, and overlapping demonstrations
abstract
Abstract In this paper, we define a novel inverse reinforcement learning (IRL) problem where the demonstrations are multi-intention, i.e., collected from multi-intention experts, unlabeled, i.e., without intention labels, and partially overlapping, i.e., shared between multiple intentions. In the presence of overlapping demonstrations, current IRL methods, developed to handle multi-intention and unlabeled demonstrations, cannot successfully learn the underlying reward functions. To solve this limitation, we propose a novel clustering-based approach to disentangle the observed demonstrations and experimentally validate its advantages. Traditional clustering-based approaches to multi-intention IRL, which are developed on the basis of model-based Reinforcement Learning (RL), formulate the problem using parametric density estimation. However, in high-dimensional environments and unknown system dynamics, i.e., model-free RL, the solution of parametric density estimation is only tractable up to the density normalization constant. To solve this, we formulate the problem as a mixture of logistic regressions to directly handle the unnormalized density. To research the challenges faced by overlapping demonstrations, we introduce the concepts of shared pair, which is a state-action pair that is shared in more than one intention, and separability, which resembles how well the multiple intentions can be separated in the joint state-action space. We provide theoretical analyses under the global optimality condition and the existence of shared pairs. Furthermore, we conduct extensive experiments on four simulated robotics tasks, extended to accept different intentions with specific levels of separability, and a synthetic driver task developed to directly control the separability. We evaluate the existing baselines on our defined problem and demonstrate, theoretically and experimentally, the advantages of our clustering-based solution, especially when the separability of the demonstrations decreases.
Ariyan Bighashdel, Pavol Jancura, Gijs Dubbelman
Mach. Learn.3
2022 Self-Supervised Road Layout Parsing with Graph Auto-Encoding
abstract
Aiming for higher-level scene understanding, this work presents a neural network approach that takes a road-layout map in bird’s-eye-view as input, and predicts a human-interpretable graph that represents the road’s topological layout. Our approach elevates the understanding of road layouts from pixel level to the level of graphs. To achieve this goal, an image-graph-image auto-encoder is utilized. The network is designed to learn to regress the graph representation at its auto-encoder bottleneck. This learning is self-supervised by an image reconstruction loss, without needing any external manual annotations. We create a synthetic dataset containing common road layout patterns and use it for training of the auto-encoder in addition to the real-world Argoverse dataset. By using this additional synthetic dataset, which conceptually captures human knowledge of road layouts and makes this available to the network for training, we are able to stabilize and further improve the performance of topological road layout understanding on the real-world Argoverse dataset. The evaluation shows that our approach exhibits comparable performance to a strong fully-supervised baseline.
Chenyang Lu 0002, Gijs Dubbelman
IV2
2022 Scene Spatio-Temporal Graph Convolutional Network for Pedestrian Intention Estimation
abstract
For safe and comfortable navigation of autonomous vehicles, it is crucial to know the pedestrian’s intention of crossing the street. Generally, human drivers are aware of the traffic objects (e.g., crosswalks and traffic lights) in the environment while driving; likewise, these objects would play a crucial role for autonomous vehicles. In this research, we propose a novel pedestrian intention estimation method that not only takes into account the influence of traffic objects but also learns their contribution levels on the intention of the pedestrian. Our proposed method, referred to as Scene SpatioTemporal Graph Convolutional Network (Scene-STGCN), takes benefits from the strength of Graph Convolutional Networks and efficiently encodes the relationships between the pedestrian and the scene objects both spatially and temporally. We conduct several experiments on the Pedestrian Intention Estimation (PIE) dataset and illustrate the importance of scene objects and their contribution levels in the task of pedestrian intention estimation. Furthermore, we perform statistical analysis on the relevance of different traffic objects in the PIE dataset and carry out an ablation study on the effect of various information sources in the scene. Finally, we demonstrate the significance of the proposed Scene-STGCN through experimental comparisons with several baselines. The results indicate that our proposed Scene-STGCN outperforms the current state-of-the-art method by 0.03 in terms of ROC-AUC metric.
Abhilash Y. Naik, Ariyan Bighashdel, Pavol Jancura, Gijs Dubbelman
IV4
2022 Learning to Predict Collision Risk from Simulated Video Data
abstract
We propose an image-based collision risk prediction model and a training strategy that allows training on simulated video data and successfully generalizes to real data. By doing so, we solve the data scarcity problem of collecting and labeling real (near) collisions, which are exceptionally rare events. Domain generalization from simulated to real data is taken into account by design by decoupling the learning strategy, and using task-specific, domain-resilient intermediate representations. Specifically, we use optical flow and vehicle bounding boxes, since they are instinctively related to the task of collision risk prediction and because their simulated-to-real domain gap is significantly lower than that of camera video data, i.e., they are more domain resilient. To demonstrate our approach, we present RiskNet, a novel neural network for image-based collision risk prediction, which classifies individual frames of a video sequence of a front-facing camera as safe or unsafe. Additionally, we present two novel datasets: the simulated Prescan dataset (which we intend to make publicly available) for training and the YouTube Driving Incidents Database (YDID) for real-world testing. The performance of RiskNet, trained solely on simulated data and tested on the real-world YDID, is comparable to that of a human driver, both in accuracy (91.8% vs. 93.6%) and F1-score (0.92 vs 0.94).
Tim J. Schoonbeek, Fabrizio J. Piva, Hamid R. Abdolhay, Gijs Dubbelman
IV4
2021 Part-Aware Panoptic Segmentation
Daan de Geus, Panagiotis Meletis, Chenyang Lu 0002, Xiaoxiao Wen, Gijs Dubbelman
CVPR5
2021 Deep Adaptive Multi-intention Inverse Reinforcement Learning
Ariyan Bighashdel, Panagiotis Meletis, Pavol Jancura, Gijs Dubbelman
ECML/PKDD (1)4
2020 A Stereo Perception Framework for Autonomous Vehicles
abstract
Stereo cameras are crucial sensors for self-driving vehicles as they are low-cost and can be used to estimate depth. It can be used for multiple purposes, such as object detection, depth estimation, semantic segmentation, etc. In this paper, we propose a stereo vision-based perception framework for autonomous vehicles. It uses three deep neural networks simultaneously to perform free-space detection, lane boundary detection, and object detection on image frames captured using the stereo camera. The depth of the detected objects from the vehicle is estimated from the disparity image computed using two stereo image frames from the stereo camera. The proposed stereo perception framework runs at 7.4 Hz on the Nvidia Drive PX 2 hardware platform, which further allows for its use in multi-sensor fusion for localization, mapping, and path planning by autonomous vehicle applications.
Narsimlu Kemsaram, Anweshan Das, Gijs Dubbelman
VTC Spring3
2020 Learning to complete partial observations from unpaired prior knowledge
abstract
We present a novel training strategy that allows convolutional encoder-decoder networks, to complete partially observed data by means of hallucination. As input, it takes data from a partially observed domain, for which no complete ground truth is available, and data from an unpaired prior knowledge domain and trains the network in an end-to-end manner. This strategy is demonstrated for the task of completing 2-D road layouts as well as 3-D vehicle shapes. In contrast to alternative approaches, our strategy is compatible with networks that use skip connections, to improve detail in the completed output, while not requiring adversarial supervision. To demonstrate its benefits, our training strategy is benchmarked against two state-of-the-art baselines, one using a two-step auto-encoder training strategy and one using an adversarial strategy. Our novel strategy achieves an improvement up to +12% F-measure on the Cityscapes dataset. The learned network intrinsically generalizes better than the baselines on unseen datasets, which is demonstrated by an improvement up to +24% F-measure on the unseen KITTI dataset. Moreover, our approach outperforms the baselines using the same backbone network on the 3-D shape completion benchmark by reducing the Hamming distance with 15%.
Chenyang Lu 0002, Gijs Dubbelman
Pattern Recognit.2
2019 Single Network Panoptic Segmentation for Street Scene Understanding
abstract
In this work, we propose a single deep neural network for panoptic segmentation, for which the goal is to provide each individual pixel of an input image with a class label, as in semantic segmentation, as well as a unique identifier for specific objects in an image, following instance segmentation. Our network makes joint semantic and instance segmentation predictions and combines these to form an output in the panoptic format. This has two main benefits: firstly, the entire panoptic prediction is made in one pass, reducing the required computation time and resources; secondly, by learning the tasks jointly, information is shared between the two tasks, thereby improving performance. Our network is evaluated on two street scene datasets: Cityscapes and Mapillary Vistas. By leveraging information exchange and improving the merging heuristics, we increase the performance of the single network, and achieve a score of 23.9 on the Panoptic Quality (PQ) metric on Mapillary Vistas validation, with an input resolution of 640 × 900 pixels. On Cityscapes validation, our method achieves a PQ score of 45.9 with an input resolution of 512 × 1024 pixels. Moreover, our method decreases the prediction time by a factor of 2 with respect to separate networks.
Daan de Geus, Panagiotis Meletis, Gijs Dubbelman
IV3
2019 On Boosting Semantic Street Scene Segmentation with Weak Supervision
abstract
Training convolutional networks for semantic segmentation requires per-pixel ground truth labels, which are very time consuming and hence costly to obtain. Therefore, in this work, we research and develop a hierarchical deep network architecture and the corresponding loss for semantic segmentation that can be trained from weak supervision, such as bounding boxes or image level labels, as well as from strong per-pixel supervision. We demonstrate that the hierarchical structure and the simultaneous training on strong (per-pixel) and weak (bounding boxes) labels, even from separate datasets, consistently increases the performance against per-pixel only training. Moreover, we explore the more challenging case of adding weak image-level labels. We collect street scene images and weak labels from the immense Open Images dataset to generate the OpenScapes dataset, and we use this novel dataset to increase segmentation performance on two established per-pixel labeled datasets, Cityscapes and Vistas. We report performance gains up to +13.2% mIoU on crucial street scene classes, and inference speed of 20 fps on a Titan V GPU for Cityscapes at 512 ×1024 resolution. Our network and OpenScapes dataset are shared with the research community.
Panagiotis Meletis, Gijs Dubbelman
IV2
2019 Incremental Hopping-Window Pose-Graph Fusion for Real-Time Vehicle Localization
abstract
In this work, we research and evaluate incremental hopping-window pose-graph fusion strategies for vehicle localization. Pose-graphs can model multiple absolute and relative vehicle localization sensors, and can be optimized using non-linear techniques. We focus on the performance of incremental hopping-window optimization for on- line usage in vehicles and compare it with global off-line optimization. Our evaluation is based on 180 Km long vehicle trajectories that are recorded in highway, urban, and rural areas, and that are accompanied with post-processed Real Time Kinematic GNSS as ground truth. The results exhibit a 17% reduction in the error's standard deviation and a significant reduction in GNSS outliers when compared with automotive-grade GNSS receivers. The incremental hopping-window pose- graph optimization bounds the computation cost, when compared to global pose-graph fusion, which increases linearly with the size of the pose- graph, whereas the difference in accuracy is only 1%. This allows real-time usage of non-linear pose-graph fusion for vehicle localization.
Anweshan Das, Gijs Dubbelman
VTC Spring2
2019 A Domain Agnostic Normalization Layer for Unsupervised Adversarial Domain Adaptation
abstract
We propose a normalization layer for unsupervised domain adaption in semantic scene segmentation. Normalization layers are known to improve convergence and generalization and are part of many state-of-the-art fully-convolutional neural networks. We show that conventional normalization layers worsen the performance of current Unsupervised Adversarial Domain Adaption (UADA), which is a method to improve network performance on unlabeled data sets and the focus of our research. Therefore, we propose a novel Domain Agnostic Normalization layer and thereby unlock the benefits of normalization layers for unsupervised adversarial domain adaptation. In our evaluation, we adapt from the synthetic GTA5 data set to the real Cityscapes data set, a common benchmark experiment, and surpass the state-of-the-art. As our normalization layer is domain agnostic at test time, we furthermore demonstrate that UADA using Domain Agnostic Normalization improves performance on unseen domains, specifically on Apolloscape and Mapillary.
Rob Romijnders, Panagiotis Meletis, Gijs Dubbelman
WACV3
2018 An Experimental Study on Relative and Absolute Pose Graph Fusion for Vehicle Localization
abstract
In this work, we research and evaluate multiple pose-graph fusion strategies for vehicle localization. We focus on fusing a single absolute localization system, i.e. automotive-grade Global Navigation Satellite System (GNSS) at 1 Hertz, with a single relative localization system, i.e. vehicle odometry at 25 Hertz. Our evaluation is based on 180 Km long vehicle trajectories that are recorded in highway, urban and rural areas, and that are accompanied with post-processed Real Time Kinematic GNSS as ground truth. The results exhibit a significant reduction in the error's standard deviation by 18% but the bias in the error is unchanged, when compared to non-fused GNSS. We show that the underlying principle is the fact that errors in GNSS readings are highly correlated in time. This causes a bias that cannot be compensated for by using the relative localization information from the odometry, but it can reduce the standard deviation of the error.
Anweshan Das, Gijs Dubbelman
Intelligent Vehicles Symposium2
2018 Training of Convolutional Networks on Multiple Heterogeneous Datasets for Street Scene Semantic Segmentation
abstract
We propose a convolutional network with hierarchical classifiers for per-pixel semantic segmentation, which is able to be trained on multiple, heterogeneous datasets and exploit their semantic hierarchy. Our network is the first to be simultaneously trained on three different datasets from the intelligent vehicles domain, i.e. Cityscapes, GTSDB and Mapillary Vistas, and is able to handle different semantic levelof-detail, class imbalances, and different annotation types, i.e. dense per-pixel and sparse bounding-box labels. We assess our hierarchical approach, by comparing against flat, nonhierarchical classifiers and we show improvements in mean pixel accuracy of 13.0% for Cityscapes classes and 2.4% for Vistas classes and 32.3% for GTSDB classes. Our implementation achieves inference rates of 17 fps at a resolution of 520 × 706 for 108 classes running on a GPU.
Panagiotis Meletis, Gijs Dubbelman
Intelligent Vehicles Symposium2
2015 COP-SLAM: Closed-Form Online Pose-Chain Optimization for Visual SLAM
abstract
In this paper, we analyze and extend the recently proposed closed-form online pose-chain simultaneous localization and mapping (SLAM) algorithm. Pose-chains are a specific type of extremely sparse pose-graphs and a product of contemporary SLAM front-ends, which perform accurate visual odometry and reliable appearance-based loop detection. They are relevant for challenging robotic applications in large-scale 3-D environments for which frequent loop detection is not desired or not possible. Closed-form online pose-chain SLAM efficiently and accurately optimizes pose-chains by exploiting their Lie group structure. The convergence and optimality properties of this solution are discussed in detail and are compared against state-of-the-art iterative methods. We also provide a novel solution space, that of similarity transforms, which has not been considered earlier for the proposed algorithm. This allows for closed-form optimization of pose-chains that exhibit scale drift, which is important to monocular SLAM systems. On the basis of extensive experiments, specifically targeting 3-D pose-chains and using a total of 60 km of challenging binocular and monocular data, it is shown that the accuracy obtained by closed-form online pose-chain SLAM is comparable with that of state-of-the-art iterative methods, while the time it needs to compute its solution is orders of magnitudes lower. This novel SLAM technique thereby is relevant to a broad range of robotic applications and computational platforms.
Gijs Dubbelman, Brett Browning
IEEE Trans. Robotics1
2014 Robust sensor cloud localization from range measurements
abstract
This work provides a feasibility study on estimating the 3-D locations of several thousand miniaturized free-floating sensor platforms. The localization is performed on basis of sparse ultrasound range measurements between sensor platforms and without the use of beacons.
Gijs Dubbelman, Erik Duisterwinke, Libertario Demi, Elena Talnishnikh, Heinrich Wörtche, Jan W. M. Bergmans
IROS1
2013 Closed-form Online Pose-chain SLAM
abstract
A novel closed-form solution for pose-graph SLAM is presented. It optimizes pose-graphs of particular structure called pose-chains by employing an extended version of trajectory bending. Our solution is designed as a back-end optimizer to be used within systems whose front-end performs state-of-the-art visual odometry and appearance based loop detection. The optimality conditions of our closed-form method and that of state-of-the-art iterative methods are discussed. The practical relevance of their theoretical differences is investigated by extensive experiments using simulated and real data. It is shown using 49 kilometers of challenging binocular data that the accuracy obtained by our closed-form solution is comparable to that of state-of-the-art iterative solutions while the time it needs to compute its solution is a factor 50 to 200 times lower. This makes our approach relevant to a broad range of applications and computational platforms.
Gijs Dubbelman, Brett Browning
ICRA1
2012 Manifold Statistics for Essential Matrices
Gijs Dubbelman, Leo Dorst, Henk Pijls
ECCV (2)1
2012 Orientation only loop-closing with closed-form trajectory bending
abstract
In earlier work closed-form trajectory bending was shown to provide an efficient and accurate out-of-core solution for loop-closing exactly sparse trajectories. Here we extend it to fuse exactly sparse trajectories, obtained from relative pose estimates, with absolute orientation data. This allows us to close-the-loop using absolute orientation data only. The benefit is that our approach does not rely on the observations from which the trajectory was estimated nor on the probabilistic links between poses in the trajectory. It therefore is highly efficient. The proposed method is compared against regular fusion and an iterative trajectory bending solution using a 5 km long urban trajectory. Proofs concerning optimality of our method are provided.
Gijs Dubbelman, Peter Hansen 0001, Brett Browning, M. Bernardine Dias
ICRA1
2012 Bias compensation in visual odometry
abstract
Empirical evidence shows that error growth in visual odometry is biased. A projective bias model is developed and its parameters are estimated offline from trajectories encompassing loops. The model is used online to compensate for bias and thereby significantly reduces error growth. We validate our approach with more than 25 km of stereo data collected in two very different urban environments from a moving vehicle. Our results demonstrate significant reduction in error, typically on the order of 50%, suggesting that our technique has significant applicability to deployed robot systems in GPS denied environments.
Gijs Dubbelman, Peter Hansen 0001, Brett Browning
IROS1
2010 Efficient trajectory bending with applications to loop closure
abstract
In robotic applications the absolute pose is often obtained as the integral of successive relative rigid-body motions. As each relative rigid-body motion is typically the product of statistical inference, the integrated absolute pose will exhibit error build-up and the estimated trajectory will differ from the true trajectory undertaken by the system. Some application areas allow the system to receive additional information about its current absolute pose, for example from loop detection, which is more accurate than the integral of the relative rigid-body motions. The availability of this absolute information is usually less frequent than the information underlying the relative rigid-body motions. This contribution addresses an efficient closed form algorithm which minimally bends a trajectory such that the integrated pose is exactly equal to any particular desired pose. The manner in which the bending is distributed over the trajectory is controllable using weights. The proposed method will be compared against a maximum likelihood solution on simulated trajectories as well as on trajectories estimated from binocular and monocular data. The results indicate that the performance differences between the closed form approach and the maximum likelihood solution are negligible while the closed form approach is significantly more efficient.
Gijs Dubbelman, Isaac Esteban, Klamer Schutte
IROS1
2009 Bias reduction for stereo based motion estimation with applications to large scale visual odometry
abstract
This contribution addresses the problem of bias in stereo based motion estimation. Using a biased estimator within a visual-odometry system will cause significant drift on large trajectories. This drift is often minimized by exploiting auxiliary sensors, (semi-)global optimization or loop-closing. In this paper it is shown that bias in the motion estimates can be caused by incorrect modeling of the uncertainties in landmark locations. Furthermore, there exists a relation between the bias, the true motion and the distribution of landmarks in space. Guided by these observations, a novel bias reduction technique has been developed. The core of the proposed method is computing the difference between motion estimates obtained using dissimilar heteroscedastic landmark uncertainty models. This approach is accurate, efficient and does not rely on auxiliary sensors, (semi-)global optimization or loop-closing. To show the real-world applicability of the proposed method, it has been tested on several data-sets including a challenging 5 km urban trajectory. The gain in performance is clearly noticeable.
Gijs Dubbelman, Frans C. A. Groen
CVPR1
2008 Accurate and robust ego-motion estimation using expectation maximization
abstract
A novel robust visual-odometry technique, called EM-SE(3) is presented and compared against using the random sample consensus (RANSAC) for ego-motion estimation. In this contribution, stereo-vision is used to generate a number of minimal-set motion hypothesis. By using EM-SE(3), which involves expectation maximization on a local linearization of the rigid-body motion group SE(3), a distinction can be made between inlier and outlier motion hypothesis. At the same time a robust mean motion as well as its associated uncertainty can be computed on the selected inlier motion hypothesis. The data-sets used for evaluation consist of synthetic and large real-world urban scenes, including several independently moving objects. Using these data-sets, it will be shown that EM-SE(3) is both more accurate and more efficient than RANSAC.
Gijs Dubbelman, Wannes van der Mark, Frans C. A. Groen
IROS1
2007 Obstacle detection during day and night conditions using stereo vision
abstract
We have developed a stereo vision based obstacle detection (OD) system that can be used to detect obstacles in off-road terrain during both day and night conditions. In order to acquire enough depth estimates for reliable OD during low visibility conditions, we propose a stereo disparity (depth) estimation approach that uses fine-to-coarse selection in a stereo image pyramid. This fine-to-coarse selection is based on a novel disparity validity metric that reflects the estimation reliability. Dense three-dimensional terrain data is reconstructed from the estimated stereo disparities. In our OD methods, several geometric properties, such as the terrain slope, are inspected to distinguish between obstacles and drivable terrain. This is achieved in a robust and efficient manner by considering the inherent uncertainty in stereo depth and using a hysteresis threshold. A large and varied collection of day- and nighttime images has been used to evaluate the performance of our system. The results show that our methods can reliably detect different types of obstacles in all tested conditions.
Gijs Dubbelman, Wannes van der Mark, Johan H. C. van den Heuvel, Frans C. A. Groen
IROS1