EDBT 2026 Demo / reviewers in the wild / expert
Claudius Gläser
dblp:63/4398
· DBLP profile ↗
22ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-9944-7405ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point CloudsabstractMasked autoencoders (MAE) have shown tremendous potential for self-supervised learning (SSL) in vision and beyond. However, point clouds from LiDARs used in automated driving are particularly challenging for MAEs since large areas of the 3D volume are empty. Consequently, existing work suffers from leaking occupancy information into the decoder and has significant computational complexity, thereby limiting the SSL pre-training to only 2D bird's eye view encoders in practice. In this work, we propose the novel neighborhood occupancy MAE (NOMAE) that overcomes the aforementioned challenges by employing masked occupancy reconstruction only in the neighborhood of non- masked voxels. We incorporate voxel masking and occupancy reconstruction at multiple scales with our proposed hierarchical mask generation technique to capture features of objects of different sizes in the point cloud. NOMAEs are extremely flexible and can be directly employed for SSL in existing 3D architectures. We perform extensive evaluations on the nuScenes and Waymo Open datasets for the downstream perception tasks of semantic segmentation and 3D object detection, comparing with both discriminative and generative SSL methods. The results demonstrate that NOMAE sets the new state-of-the-art on multiple benchmarks for multiple point cloud perception tasks. Mohamed Abdelsamad, Michael Ulrich, Claudius Gläser, Abhinav Valada |
CVPR | 3 |
| 2025 | Open-Set LiDAR Panoptic Segmentation Guided by Uncertainty-Aware LearningabstractAutonomous vehicles that navigate in open-world environments may encounter previously unseen object classes. However, most existing LiDAR panoptic segmentation models rely on closed-set assumptions, failing to detect unknown object instances. In this work, we propose ULOPS, an uncertainty-guided open-set panoptic segmentation framework that leverages Dirichlet-based evidential learning to model predictive uncertainty. Our architecture incorporates separate decoders for semantic segmentation with uncertainty estimation, embedding with prototype association, and instance center prediction. During inference, we leverage uncertainty estimates to identify and segment unknown instances. To strengthen the model’s ability to differentiate between known and unknown objects, we introduce three uncertainty-driven loss functions. Uniform Evidence Loss to encourage high uncertainty in unknown regions. Adaptive Uncertainty Separation Loss ensures a consistent difference in uncertainty estimates between known and unknown objects at a global scale. Contrastive Uncertainty Loss refines this separation at the fine-grained level. To evaluate open-set performance, we extend benchmark settings on KITTI-360 and introduce a new open-set evaluation for nuScenes. Extensive experiments demonstrate that ULOPS consistently outperforms existing open-set LiDAR panoptic segmentation methods. We make the code and pre-trained models available at http://ulops.cs.uni-freiburg.de. Rohit Mohan, Julia Hindel, Florian Drews, Claudius Gläser, Daniele Cattaneo 0001, Abhinav Valada |
IROS | 4 |
| 2024 | Revisiting Out-of-Distribution Detection in LiDAR-based 3D Object DetectionabstractLiDAR-based 3D object detection has become an essential part of automated driving due to its ability to localize and classify objects precisely in 3D. However, object detectors face a critical challenge when dealing with unknown foreground objects, particularly those that were not present in their original training data. These out-of-distribution (OOD) objects can lead to misclassifications, posing a significant risk to the safety and reliability of automated vehicles. Currently, LiDAR-based OOD object detection has not been well studied. We address this problem by generating synthetic training data for OOD objects by perturbing known object categories. Our idea is that these synthetic OOD objects produce different responses in the feature map of an object detector compared to in-distribution (ID) objects. We then extract features using a pre-trained and fixed object detector and train a simple multilayer perceptron (MLP) to classify each detection as either ID or OOD. In addition, we propose a new evaluation protocol that allows the use of existing datasets without modifying the point cloud, ensuring a more authentic evaluation of real-world scenarios. The effectiveness of our method is validated through experiments on the newly proposed nuScenes OOD benchmark. The source code is available at https://github.com/uulm-mrm/mmood3d. Michael Kösel, Marcel Schreiber, Michael Ulrich, Claudius Gläser, Klaus Dietmayer |
IV | 4 |
| 2023 | Self-Supervised Occupancy Grid Map Completion for Automated DrivingabstractThis paper investigates methods for enhancing the quality of occupancy grid maps (OGMs) using a combination of a self-supervised data generation procedure using only unlabeled data and a deep learning approach. OGMs are grid-structured environment representations, commonly used in automated driving systems to encode occupancy of the surrounding area. However, due to limited sensor range and resolution, their quality degrades significantly in distant and occluded areas, posing a challenge for a subsequent decision making. We introduce OGM completion, whose goal is to provide a more complete representation of the environment by extrapolating potential occupancy to distant and occluded areas. In particular, we propose and implement a complete framework for OGM completion. We develop a method for self-supervised data generation, identify an existing class of adoptable deep learning architectures, adapt loss functions and a quantitative performance metric, and derive a generic baseline method. Finally, we validate the functionality of the implemented framework by thorough experimentation and inspection of real-world examples of OGM completion in automated driving, significantly outperforming a baseline method. Jugoslav Stojcheski, Thomas Nürnberg, Michael Ulrich, Claudius Gläser, Andreas Geiger 0001 |
IV | 5 |
| 2022 | DeepFusion: A Robust and Modular 3D Object Detector for Lidars, Cameras and RadarsabstractWe propose DeepFusion, a modular multi-modal architecture to fuse lidars, cameras and radars in different combinations for 3D object detection. Specialized feature extractors take advantage of each modality and can be exchanged easily, making the approach simple and flexible. Extracted features are transformed into bird's-eye-view as a common representation for fusion. Spatial and semantic alignment is performed prior to fusing modalities in the feature space. Finally, a detection head exploits rich multi-modal features for improved 3D detection performance. Experimental results for lidar-camera, lidar-camera-radar and camera-radar fusion show the flexibility and effectiveness of our fusion approach. In the process, we study the largely unexplored task of faraway car detection up to 225 meters, showing the benefits of our lidar-camera fusion. Furthermore, we investigate the required density of lidar points for 3D object detection and illustrate implications at the example of robustness against adverse weather conditions. Moreover, ablation studies on our camera-radar fusion highlight the importance of accurate depth estimation. Florian Drews, Di Feng, Florian Faion, Lars Rosenbaum, Michael Ulrich, Claudius Gläser |
IROS | 6 |
| 2022 | Self-Supervised Velocity Estimation for Automotive Radar Object Detection NetworksabstractThis paper presents a method to learn the Cartesian velocity of objects using an object detection network on automotive radar data. The proposed method is self-supervised in terms of generating its own training signal for the velocities. Labels are only required for single-frame, oriented bounding boxes (OBBs). Labels for the Cartesian velocities or contiguous sequences, which are expensive to obtain, are not required. The general idea is to pre-train an object detection network without velocities using single-frame OBB labels, and then exploit the network’s OBB predictions on unlabelled data for velocity training. In detail, the network’s OBB predictions of the unlabelled frames are updated to the timestamp of a labelled frame using the predicted velocities and the distances between the updated OBBs of the unlabelled frame and the OBB predictions of the labelled frame are used to generate a self-supervised training signal for the velocities. The detection network architecture is extended by a module to account for the temporal relation of multiple scans and a module to represent the radars’ radial velocity measurements explicitly. A twostep approach of first training only OBB detection, followed by training OBB detection and velocities is used. Further, a pre-training with pseudo-labels generated from radar radial velocity measurements bootstraps the self-supervised method of this paper. Experiments on the publicly available nuScenes dataset show that the proposed method almost reaches the velocity estimation performance of a fully supervised training, but does not require expensive velocity labels. Furthermore, we outperform a baseline method which uses only radial velocity measurements as labels. Daniel Niederlöhner, Michael Ulrich, Sascha Braun, Daniel Köhler, Florian Faion, Claudius Gläser, André Treptow, Holger Blume |
IV | 6 |
| 2022 | Transformers for Multi-Object Tracking on Point CloudsabstractWe present TransMOT, a novel transformer-based end-to-end trainable online tracker and detector for point cloud data. The model utilizes a cross- and a self-attention mechanism and is applicable to lidar data in an automotive context, as well as other data types, such as radar. Both track management and the detection of new tracks are performed by the same transformer decoder module and the tracker state is encoded in feature space. With this approach, we make use of the rich latent space of the detector for tracking rather than relying on low-dimensional bounding boxes. Still, we are able to retain some of the desirable properties of traditional Kalman-filter based approaches, such as an ability to handle sensor input at arbitrary timesteps or to compensate frame skips. This is possible due to a novel module that transforms the track information from one frame to the next on feature-level and thereby fulfills a similar task as the prediction step of a Kalman filter. Results are presented on the challenging real-world dataset nuScenes, where the proposed model outperforms its Kalman filter-based tracking baseline. Felicia Ruppel, Florian Faion, Claudius Gläser, Klaus Dietmayer |
IV | 3 |
| 2022 | A Multi-Task Recurrent Neural Network for End-to-End Dynamic Occupancy Grid MappingabstractA common approach for modeling the environment of an autonomous vehicle are dynamic occupancy grid maps, in which the surrounding is divided into cells, each containing the occupancy and velocity state of its location. Despite the advantage of modeling arbitrary shaped objects, the used algorithms rely on hand-designed inverse sensor models and semantic information is missing. Therefore, we introduce a multi-task recurrent neural network to predict grid maps providing occupancies, velocity estimates, semantic information and the driveable area. During training, our network architecture, which is a combination of convolutional and recurrent layers, processes sequences of raw lidar data, that is represented as bird’s eye view images with several height channels. The multi-task network is trained in an end-to-end fashion to predict occupancy grid maps without the usual preprocessing steps consisting of removing ground points and applying an inverse sensor model. In our evaluations, we show that our learned inverse sensor model is able to overcome some limitations of a geometric inverse sensor model in terms of representing object shapes and modeling freespace. Moreover, we report a better runtime performance and more accurate semantic predictions for our end-to-end approach, compared to our network relying on measurement grid maps as input data. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
IV | 3 |
| 2022 | Pedestrian Behavior Prediction for Automated Driving: Requirements, Metrics, and Relevant FeaturesabstractAutomated vehicles require a comprehensive understanding of traffic situations to ensure safe and anticipatory driving. In this context, the prediction of pedestrians is particularly challenging as pedestrian behavior can be influenced by multiple factors. In this paper, we thoroughly analyze the requirements on pedestrian behavior prediction for automated driving via a system-level approach. To this end we investigate real-world pedestrian-vehicle interactions with human drivers. Based on human driving behavior we then derive appropriate reaction patterns of an automated vehicle and determine requirements for the prediction of pedestrians. This includes a novel metric tailored to measure prediction performance from a system-level perspective. The proposed metric is evaluated on a large-scale dataset comprising thousands of real-world pedestrian-vehicle interactions. We furthermore conduct an ablation study to evaluate the importance of different contextual cues and compare these results to ones obtained using established performance metrics for pedestrian prediction. Our results highlight the importance of a system-level approach to pedestrian behavior prediction. Michael Herman, Vishnu Suganth Prabhakaran, Nicolas Möser, Hanna Carolin Ziesche, Waleed Ahmed, Lutz Bürkle, Ernst Kloppenburg, Claudius Gläser |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2021 | Dynamic Occupancy Grid Mapping with Recurrent Neural NetworksabstractModeling and understanding the environment is an essential task for autonomous driving. In addition to the detection of objects, in complex traffic scenarios the motion of other road participants is of special interest. Therefore, we propose to use a recurrent neural network to predict a dynamic occupancy grid map, which divides the vehicle surrounding in cells, each containing the occupancy probability and a velocity estimate. During training, our network is fed with sequences of measurement grid maps, which encode the lidar measurements of a single time step. Due to the combination of convolutional and recurrent layers, our approach is capable to use spatial and temporal information for the robust detection of static and dynamic environment. In order to apply our approach with measurements from a moving ego-vehicle, we propose a method for ego-motion compensation that is applicable in neural network architectures with recurrent layers working on different resolutions. In our evaluations, we compare our approach with a state-of-the-art particle-based algorithm on a large publicly available dataset to demonstrate the improved accuracy of velocity estimates and the more robust separation of the environment in static and dynamic area. Additionally, we show that our proposed method for ego-motion compensation leads to comparable results in scenarios with stationary and with moving ego-vehicle. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
ICRA | 3 |
| 2021 | Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and ChallengesabstractRecent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs, Radars), and multiple sensing modalities can be fused to exploit their complementary properties. In this context, many methods have been proposed for deep multi-modal perception problems. However, there is no general guideline for network architecture design, and questions of “what to fuse”, “when to fuse”, and “how to fuse” remain open. This review paper attempts to systematically summarize methodologies and discuss challenges for deep multi-modal object detection and semantic segmentation in autonomous driving. To this end, we first provide an overview of on-board sensors on test vehicles, open datasets, and background information for object detection and semantic segmentation in autonomous driving research. We then summarize the fusion methodologies and discuss challenges and open questions. In the appendix, we provide tables that summarize topics and methods. We also provide an interactive online platform to navigate each reference: https://boschresearch.github.io/multimodalperception/. Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Gläser, Fabian Timm, Werner Wiesbeck, Klaus Dietmayer |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Motion Estimation in Occupancy Grid Maps in Stationary Settings Using Recurrent Neural NetworksabstractIn this work, we tackle the problem of modeling the vehicle environment as dynamic occupancy grid map in complex urban scenarios using recurrent neural networks. Dynamic occupancy grid maps represent the scene in a bird's eye view, where each grid cell contains the occupancy probability and the two dimensional velocity. As input data, our approach relies on measurement grid maps, which contain occupancy probabilities, generated with lidar measurements. Given this configuration, we propose a recurrent neural network architecture to predict a dynamic occupancy grid map, i.e. filtered occupancy and velocity of each cell, by using a sequence of measurement grid maps. Our network architecture contains convolutional long-short term memories in order to sequentially process the input, makes use of spatial context, and captures motion. In the evaluation, we quantify improvements in estimating the velocity of braking and turning vehicles compared to the state-of-the-art. Additionally, we demonstrate that our approach provides more consistent velocity estimates for dynamic objects, as well as, less erroneous velocity estimates in static area. Marcel Schreiber, Vasileios Belagiannis, Claudius Gläser, Klaus Dietmayer |
ICRA | 3 |
| 2016 | The narrow road assistant - evolution towards highly automated driving in inner cityabstractWith the “narrow road assistant” (NRA), this contribution elaborates on an inner-city ADAS that supports the driver in safely passing narrow road passages. The system offers early scenario-specific information together with automated support in steering and braking. The solved challenges and gathered experiences will ease and accelerate the development of future functions for piloted inner-city driving. In this sense, the function is an important evolutionary step towards highly automated driving in urban areas. The ADAS was realized in the context of the publicly-funded project UR:BAN, which started in 2012. In October 2015 the project UR:BAN reached its peak in a public presentation in Düsseldorf Germany. Numerous driving demonstrations allowed guests an independent evaluation on a larger scale. The Bosch NRA performed in all situations reliably surpassing the expectations of many guests. The focus of the contribution will be on the detailed description of the system and its evaluation. Claudius Gläser, Lutz Bürkle, Frank Niewels |
Intelligent Vehicles Symposium | 2 |
| 2014 | Environment perception for inner-city driver assistance and highly-automated drivingabstractWhile driver assistance systems mainly targeted highway or parking scenarios in the past, systems assisting in inner-city traffic increasingly get into focus today. Since driving in urban areas is characterized by a larger variety of situations that have to be covered, finding an adequate representation of the vehicle surrounding is a challenging task for the perception of these systems. In this paper we present our perception system that has been specifically designed for the demands of inner-city driving. It first features a plugin-based architecture by which multiple sensor setups as well as different driver assistance functions can be supported. Second, it is characterized by a hybrid modeling approach that combines the well known model-based object tracking technique with model-free representations in terms of grids. We will present details on a specific implementation of the system using a 3D lidar sensor. Finally, it is shown how the system is used in the Narrow Road Assistant — a next-generation driver assistance system supporting the driver in safely passing narrow road passages in inner-city. Claudius Gläser, Lutz Bürkle, Frank Niewels |
Intelligent Vehicles Symposium | 1 |
| 2011 | Discriminant Sub-Space Projection of Spectro-Temporal Speech Features Based on Maximizing Mutual InformationabstractWe previously developed noise robust Hierarchical Spectro-Temporal (HIST) speech features. The learning of the features was performed in an unsupervised way with unlabeled speech data. In a final stage we deployed Principal Component Anal-ysis (PCA) to reduce the feature dimensions and to diagonalize them. In this paper we investigate if a discriminant projection can further increase the performance. We maximize the mu-tual information between the features and the phoneme cate-gories using a procedure known as Maximizing Renyi’s Mu-tual Information (MRMI) and also compare it to Linear Dis-criminant Analysis (LDA). Based on recognition tests in clean and in noise, i. e. in matching and mismatching conditions, we show that the discriminant projections increases recogni-tion scores compared to PCA in matching conditions. How-ever, this improvement does not transfer to the mismatching, i. e. noisy, conditions. We discuss measures to alleviate this problem. Overall MRMI performs better than LDA. Index Terms: Spectro-temporal, discriminant, mutual informa-tion, robust speech recognition, auditory Martin Heckmann, Claudius Gläser |
INTERSPEECH | 2 |
| 2010 | An adaptive Normalized Gaussian Network and its application to online category learningabstractIn online applications, where training samples sequentially arise during execution, incremental learning schemes have to be applied. In this paper we propose an adaptive Normalized Gaussian Network model (NGnet) suitable for incremental learning. Following a statistical account we present a truly sequential training procedure. Key to the learning algorithm are local unit manipulation mechanisms for network growth and pruning which continuously adapt the network's complexity according to task demands. We evaluate our model in artificial and real-world categorization tasks. Thereby, we additionally introduce a framework for the categorization on adaptive feature spaces. In the system, a simultaneous extraction of class-discriminative features facilitates the NGnet's categorization of input patterns. We present simulation results which demonstrate that the framework realizes a rapid learning from few examples, small-sized network models, and an improved generalization ability. A comparison to incremental support vector machine classification yields a favorable performance of our model. Claudius Gläser, Frank Joublin |
IJCNN | 1 |
| 2010 | Applying geometric source separation for improved pitch extraction in human-robot interaction
Martin Heckmann, Claudius Gläser, Frank Joublin, Kazuhiro Nakadai |
INTERSPEECH | 2 |
| 2010 | Combining Auditory Preprocessing and Bayesian Estimation for Robust Formant TrackingabstractWe present a framework for estimating formant trajectories. Its focus is to achieve high robustness in noisy environments. Our approach combines a preprocessing based on functional principles of the human auditory system and a probabilistic tracking scheme. For enhancing the formant structure in spectrograms we use a Gammatone filterbank, a spectral preemphasis, as well as a spectral filtering using difference-of-Gaussians (DoG) operators. Finally, a contrast enhancement mimicking a competition between filter responses is applied. The probabilistic tracking scheme adopts the mixture modeling technique for estimating the joint distribution of formants. In conjunction with an algorithm for adaptive frequency range segmentation as well as Bayesian smoothing an efficient framework for estimating formant trajectories is derived. Comprehensive evaluations of our method on the VTR-formant database emphasize its high precision and robustness. We obtained superior performance compared to existing approaches for clean as well as echoic noisy speech. Finally, an implementation of the framework within the scope of an online system using instantaneous feature-based resynthesis demonstrates its applicability to real-world scenarios. Claudius Gläser, Martin Heckmann, Frank Joublin, Christian Goerick |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Enhancing Topology Preservation during Neural Field Development Via Wiring Length Minimization
Claudius Gläser, Frank Joublin, Christian Goerick |
ICANN (1) | 1 |
| 2008 | Auditory-based formant estimation in noise using a probabilistic frameworkabstractWe recently introduced a computationally efficient framework for tracking formants which combines a biologically inspired preprocessing for enhancing formants in spectrograms with a probabilistic framework for estimating formant trajectories. In contrast to previously published approaches our tracking scheme relies on the joint distribution of formants rather than using independent tracking instances for each formant separately. Therewith more precise formant estimates could be obtained. In this paper we will briefly review our algorithm and extend it by using more sophisticated models of the formants underlying dynamics. Furthermore, we will dwell on the robustness of our method for speech degraded by various types of noise. A comprehensive evaluation on a large publicly available database containing hand-labeled formant trajectories shows significant performance improvements in both clean and noisy speech compared to state of the art approaches. Claudius Gläser, Martin Heckmann, Frank Joublin, Christian Goerick |
INTERSPEECH | 1 |
| 2008 | Listen to the parrot: Demonstrating the quality of online pitch and formant extraction via feature-based resynthesisabstractWe present a system for online extraction of the fundamental frequency and the first four formant frequencies from a speech signal. In order to evaluate the performance of the extraction a resynthesis of the speech signal is performed. The resynthesis is based on the extracted frequencies and the energy of the input signal at the formant locations. The extraction of the fundamental frequency and the formants is robust against room echoes and interfering noise. In order to improve the robustness against background noise a noise reduction was implemented. Tests in three rooms of different size at varying distances to the system (up to 8 m yielding an SNR of approx. 0 dB) were performed. Martin Heckmann, Claudius Gläser, Miguel Vaz, Tobias Rodemann, Frank Joublin, Christian Goerick |
IROS | 2 |
| 2007 | Joint Estimation of Formant Trajectories via Spectro-Temporal Smoothing and Bayesian TechniquesabstractWe propose a method for the joint estimation of formant trajectories from spectrograms. Formants are enhanced in the spectrograms obtained from the application of a Gammatone filterbank via a smoothing along the frequency axis. In contrast to previously published approaches, the used tracking algorithm relies on the joint distribution of formants rather than using independent tracker instances. More precisely, Bayesian mixture filtering in conjunction with adaptive frequency range segmentation as well as Bayesian smoothing are used. The algorithm was evaluated on a publicly available database containing hand-labeled formant tracks. Experimental results show a significant performance improvement compared to a state of the art approach. Claudius Gläser, Martin Heckmann, Frank Joublin, Christian Goerick, Horst-Michael Groß |
ICASSP (4) | 1 |