VLDB 2026 Research / reviewers in the wild / expert
Oliver Wasenmüller
dblp:166/4611
· DBLP profile ↗
26ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 17 · 8 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DualSight: Learning to Disentangle Artifact and Semantic Features for Detection of Diffusion-Generated Images
Nikolas Ebert, Oliver Wasenmüller |
ICPR (1) | 3 |
| 2026 | Efficient Post-hoc Calibration in Object Detection Without Held-Out Data
Nikolas Ebert, Didier Stricker, Oliver Wasenmüller |
ICPR (1) | 3 |
| 2026 | Revisiting Point Cloud Representations Across Heterogeneous Sensors
Laurenz Reichardt, Mario Speckert, Alexander Musiat, Oliver Wasenmüller |
ICPR (1) | 4 |
| 2024 | Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis
Nikolas Ebert, Oliver Wasenmüller |
ACCV (7) | 3 |
| 2024 | GenFormer - Generated Images Are All You Need to Improve Robustness of Transformers on Small Datasets
Sven Oehri, Nikolas Ebert, Didier Stricker, Oliver Wasenmüller |
ICPR (2) | 5 |
| 2024 | Text3DAug - Prompted Instance Augmentation for LiDAR PerceptionabstractLiDAR data of urban scenarios poses unique challenges, such as heterogeneous characteristics and inherent class imbalance. Therefore, large-scale datasets are necessary to apply deep learning methods. Instance augmentation has emerged as an efficient method to increase dataset diversity. However, current methods require the time-consuming curation of 3D models or costly manual data annotation. To overcome these limitations, we propose Text3DAug, a novel approach leveraging generative models for instance augmentation. Text3DAug does not depend on labeled data and is the first of its kind to generate instances and annotations from text. This allows for a fully automated pipeline, eliminating the need for manual effort in practical applications. Additionally, Text3DAug is sensor agnostic and can be applied regardless of the LiDAR sensor used. Comprehensive experimental analysis on LiDAR segmentation, detection and novel class discovery demonstrates that Text3DAug is effective in supplementing existing methods or as a standalone method, performing on par or better than established methods, however while overcoming their specific drawbacks. The code is publicly available.1 Laurenz Reichardt, Luca Uhr, Oliver Wasenmüller |
IROS | 3 |
| 2022 | Multitask Network for Joint Object Detection, Semantic Segmentation and Human Pose Estimation in Vehicle Occupancy MonitoringabstractIn order to ensure safe autonomous driving, precise information about the conditions in and around the vehicle must be available. Accordingly, the monitoring of occupants and objects inside the vehicle is crucial. In the state-of-the-art, single or multiple deep neural networks are used for either object recognition, semantic segmentation, or human pose estimation. In contrast, we propose our Multitask Detection, Segmentation and Pose Estimation Network (MDSP) – the first multitask network solving all these three tasks jointly in the area of occupancy monitoring. Due to the shared architecture, memory and computing costs can be saved while achieving higher accuracy. Furthermore, our architecture allows a flexible combination of the three mentioned tasks during a simple end-to-end training. We perform comprehensive evaluations on the public datasets SVIRO and TiCaM in order to demonstrate the superior performance. Nikolas Ebert, Patrick Mangat, Oliver Wasenmüller |
IV | 3 |
| 2021 | Autoencoder Based Inter-Vehicle Generalization for In-Cabin Occupant ClassificationabstractCommon domain shift problem formulations consider the integration of multiple source domains, or the target domain during training. Regarding the generalization of machine learning models between different car interiors, we formulate the criterion of training in a single vehicle: without access to the target distribution of the vehicle the model would be deployed to, neither with access to multiple vehicles during training. We performed an investigation on the SVIRO dataset for occupant classification on the rear bench and propose an autoencoder based approach to improve the transferability. The autoencoder is on par with commonly used classification models when trained from scratch and sometimes out-performs models pre-trained on a large amount of data. Moreover, the autoencoder can transform images from unknown vehicles into the vehicle it was trained on. These results are corroborated by an evaluation on real infrared images from two vehicle interiors. Steve Dias Da Cruz, Bertram Taetz, Oliver Wasenmüller, Thomas Stifter, Didier Stricker |
IV | 3 |
| 2021 | SSGP: Sparse Spatial Guided Propagation for Robust and Generic InterpolationabstractInterpolation of sparse pixel information towards a dense target resolution finds its application across multiple disciplines in computer vision. State-of-the-art interpolation of motion fields applies model-based interpolation that makes use of edge information extracted from the target image. For depth completion, data-driven learning approaches are widespread. Our work is inspired by latest trends in depth completion that tackle the problem of dense guidance for sparse information. We extend these ideas and create a generic cross-domain architecture that can be applied for a multitude of interpolation problems like optical flow, scene flow, or depth completion. In our experiments, we show that our proposed concept of Sparse Spatial Guided Propagation (SSGP) achieves improvements to robustness, accuracy, or speed compared to specialized algorithms. René Schuster, Oliver Wasenmüller, Christian Unger, Didier Stricker |
WACV | 2 |
| 2020 | Ghost Target Detection in 3D Radar Data using Point Cloud based Deep Neural NetworkabstractGhost targets are targets that appear at wrong locations in radar data and are caused by the presence of multiple indirect reflections between the target and the sensor. In this work, we introduce the first point based deep learning approach for ghost target detection in 3D radar point clouds. This is done by extending the PointNet network architecture by modifying its input to include radar point features beyond location and introducing skip connetions. We compare different input modalities and analyze the effects of the changes we introduced. We also propose an approach for automatic labeling of ghost targets 3D radar data using lidar as reference. The algorithm is trained and tested on real data in various driving scenarios and the tests show promising results in classifying real and ghost radar targets. Mahdi Chamseddine, Jason R. Rambach, Didier Stricker, Oliver Wasenmüller |
ICPR | 4 |
| 2020 | HPERL: 3D Human Pose Estimation from RGB and LiDARabstractIn-the-wild human pose estimation has a huge potential for various fields, ranging from animation and action recognition to intention recognition and prediction for autonomous driving. The current state-of-the-art is focused only on RGB and RGB-D approaches for predicting the 3D human pose. However, not using precise LiDAR depth information limits the performance and leads to very inaccurate absolute pose estimation. With LiDAR sensors becoming more affordable and common on robots and autonomous vehicle setups, we propose an end-to-end architecture using RGB and LiDAR to predict the absolute 3D human pose with unprecedented precision. Additionally, we introduce a weakly-supervised approach to generate 3D predictions using 2D pose annotations from PedX [1]. This allows for many new opportunities in the field of 3D human pose estimation. Michael Fürst, Shriya T. P. Gupta, René Schuster, Oliver Wasenmüller, Didier Stricker |
ICPR | 4 |
| 2020 | ResFPN: Residual Skip Connections in Multi-Resolution Feature Pyramid Networks for Accurate Dense Pixel MatchingabstractDense pixel matching is required for many computer vision algorithms such as disparity, optical flow or scene flow estimation. Feature Pyramid Networks (FPN) have proven to be a suitable feature extractor for CNN-based dense matching tasks. FPN generates well localized and semantically strong features at multiple scales. However, the generic FPN is not utilizing its full potential, due to its reasonable but limited localization accuracy. Thus, we present ResFPN - a multi-resolution feature pyramid network with multiple residual skip connections, where at any scale, we leverage the information from higher resolution maps for stronger and better localized features. In our ablation study, we demonstrate the effectiveness of our novel architecture with clearly higher accuracy than FPN. In addition, we verify the superior accuracy of ResFPN in many different pixel matching applications on established datasets like KITTI, Sintel, and FlyingThings3D. Rishav, René Schuster, Ramy Battrawy, Oliver Wasenmüller, Didier Stricker |
ICPR | 4 |
| 2020 | DeepLiDARFlow: A Deep Learning Architecture For Scene Flow Estimation Using Monocular Camera and Sparse LiDARabstractScene flow is the dense 3D reconstruction of motion and geometry of a scene. Most state-of-the-art methods use a pair of stereo images as input for full scene reconstruction. These methods depend a lot on the quality of the RGB images and perform poorly in regions with reflective objects, shadows, ill-conditioned light environment and so on. LiDAR measurements are much less sensitive to the aforementioned conditions but LiDAR features are in general unsuitable for matching tasks due to their sparse nature. Hence, using both LiDAR and RGB can potentially overcome the individual disadvantages of each sensor by mutual improvement and yield robust features which can improve the matching process. In this paper, we present DeepLiDARFlow, a novel deep learning architecture which fuses high level RGB and LiDAR features at multiple scales in a monocular setup to predict dense scene flow. Its performance is much better in the critical regions where image-only and LiDAR-only methods are inaccurate. We verify our DeepLiDARFlow using the established data sets KITTI and FlyingThings3D and we show strong robustness compared to several state-of-the-art methods which used other input modalities. The code of our paper is available at https://github.com/dfki-av/DeepLiDARFlow. Rishav, Ramy Battrawy, René Schuster, Oliver Wasenmüller, Didier Stricker |
IROS | 4 |
| 2020 | SVIRO: Synthetic Vehicle Interior Rear Seat Occupancy Dataset and BenchmarkabstractWe release SVIRO, a synthetic dataset for sceneries in the passenger compartment of ten different vehicles, in order to analyze machine learning-based approaches for their generalization capacities and reliability when trained on a limited number of variations (e.g. identical backgrounds and textures, few instances per class). This is in contrast to the intrinsically high variability of common benchmark datasets, which focus on improving the state-of-the-art of general tasks. Our dataset contains bounding boxes for object detection, instance segmentation masks, keypoints for pose estimation and depth images for each synthetic scenery as well as images for each individual seat for classification. The advantage of our use-case is twofold: The proximity to a realistic application to benchmark new approaches under novel circumstances while reducing the complexity to a more tractable environment, such that applications and theoretical questions can be tested on a more challenging dataset as toy problems. The data and evaluation server are available under https://sviro.kl.dfki.de. Steve Dias Da Cruz, Oliver Wasenmüller, Hans-Peter Beise, Thomas Stifter, Didier Stricker |
WACV | 2 |
| 2020 | SceneFlowFields++: Multi-frame Matching, Visibility Prediction, and Robust Interpolation for Scene Flow Estimation
René Schuster, Oliver Wasenmüller, Christian Unger, Georg Kuschk, Didier Stricker |
Int. J. Comput. Vis. | 2 |
| 2019 | A Compact Light Field Camera for Real-Time Depth Estimation
Yuriy Anisimov, Oliver Wasenmüller, Didier Stricker |
CAIP (1) | 2 |
| 2019 | SDC - Stacked Dilated Convolution: A Unified Descriptor Network for Dense Matching TasksabstractDense pixel matching is important for many computer vision tasks such as disparity and flow estimation. We present a robust, unified descriptor network that considers a large context region with high spatial variance. Our network has a very large receptive field and avoids striding layers to maintain spatial resolution. These properties are achieved by creating a novel neural network layer that consists of multiple, parallel, stacked dilated convolutions (SDC). Several of these layers are combined to form our SDC descriptor network. In our experiments, we show that our SDC features outperform state-of-the-art feature descriptors in terms of accuracy and robustness. In addition, we demonstrate the superior performance of SDC in state-of-the-art stereo matching, optical flow and scene flow algorithms on several famous public benchmarks. René Schuster, Oliver Wasenmüller, Christian Unger, Didier Stricker |
CVPR | 2 |
| 2019 | LiDAR-Flow: Dense Scene Flow Estimation from Sparse LiDAR and Stereo ImagesabstractWe propose a new approach called LiDAR-Flow to robustly estimate a dense scene flow by fusing a sparse LiDAR with stereo images. We take the advantage of the high accuracy of LiDAR to resolve the lack of information in some regions of stereo images due to textureless objects, shadows, ill-conditioned light environment and many more. Additionally, this fusion can overcome the difficulty of matching unstructured 3D points between LiDAR-only scans. Our LiDAR-Flow approach consists of three main steps; each of them exploits LiDAR measurements. First, we build strong seeds from LiDAR to enhance the robustness of matches between stereo images. The imagery part seeks the motion matches and increases the density of scene flow estimation. Then, a consistency check employs LiDAR seeds to remove the possible mismatches. Finally, LiDAR measurements constraint the edge-preserving interpolation method to fill the remaining gaps. In our evaluation we investigate the individual processing steps of our LiDAR-Flow approach and demonstrate the superior performance compared to image-only approach. Ramy Battrawy, René Schuster, Oliver Wasenmüller, Qing Rao, Didier Stricker |
IROS | 3 |
| 2019 | PWOC-3D: Deep Occlusion-Aware End-to-End Scene Flow EstimationabstractIn the last few years, convolutional neural networks (CNNs) have demonstrated increasing success at learning many computer vision tasks including dense estimation problems such as optical flow and stereo matching. However, the joint prediction of these tasks, called scene flow, has traditionally been tackled using slow classical methods based on primitive assumptions which fail to generalize. The work presented in this paper overcomes these drawbacks efficiently (in terms of speed and accuracy) by proposing PWOC-3D, a compact CNN architecture to predict scene flow from stereo image sequences in an end-to-end supervised setting. Further, large motion and occlusions are well-known problems in scene flow estimation. PWOC-3D employs specialized design decisions to explicitly model these challenges. In this regard, we propose a novel self-supervised strategy to predict occlusions from images (learned without any labeled occlusion data). Leveraging several such constructs, our network achieves competitive results on the KITTI benchmark and the challenging FlyingThings3D dataset. Especially on KITTI, PWOC-3D achieves the second place among end-to-end deep learning methods with 48 times fewer parameters than the top-performing method. Rohan Saxena, René Schuster, Oliver Wasenmüller, Didier Stricker |
IV | 3 |
| 2019 | DeLiO: Decoupled LiDAR OdometryabstractMost LiDAR odometry algorithms estimate the transformation between two consecutive frames by estimating the rotation and translation in an intervening fashion. In this paper, we propose our Decoupled LiDAR Odometry (DeLiO), which - for the first time - decouples the rotation estimation completely from the translation estimation. In particular, the rotation is estimated by extracting the surface normals from the input point clouds and tracking their characteristic pattern on a unit sphere. Using this rotation the point clouds are unrotated so that the underlying transformation is pure translation, which can be easily estimated using a line cloud approach. An evaluation is performed on the KITTI dataset and the results are compared against state-of-the-art algorithms. Queens Maria Thomas, Oliver Wasenmüller, Didier Stricker |
IV | 2 |
| 2018 | FlowFields++: Accurate Optical Flow Correspondences Meet Robust InterpolationabstractOptical Flow algorithms are of high importance for many applications. Recently, the Flow Field algorithm and its modifications have shown remarkable results, as they have been evaluated with top accuracy on different data sets. In our analysis of the algorithm we have found that it produces accurate sparse matches, but there is room for improvement in the interpolation. Thus, we propose in this paper FlowFields++, where we combine the accurate matches of Flow Fields with a robust interpolation. In addition, we propose improved variational optimization as post-processing. Our new algorithm is evaluated on the challenging KITTI and MPI Sintel data sets with public top results on both benchmarks. René Schuster, Christian Bailer, Oliver Wasenmüller, Didier Stricker |
ICIP | 3 |
| 2018 | Improving Time-of-Flight Sensor for Specular Surfaces with Shape from PolarizationabstractTime-of-Flight (ToF) sensors can obtain depth values for diffuse objects. However, the essential problem is that the sensor can not receive active light from specular surfaces due to specular reflections. In this paper, we propose a new depth reconstruction framework for specular objects that combines ToF cues and Shape from Polarization (SfP). To overcome the ill-posedness of SfP with a single view, we integrate superpixel segmentation with planarity constraints for every superpixel. Experimental results demonstrate the effectiveness of the depth reconstruction algorithm for both controlled environment data and real vehicle data in a parking area. Tomonari Yoshida, Vladislav Golyanik, Oliver Wasenmüller, Didier Stricker |
ICIP | 3 |
| 2018 | SceneFlowFields: Dense Interpolation of Sparse Scene Flow CorrespondencesabstractWhile most scene flow methods use either variational optimization or a strong rigid motion assumption, we show for the first time that scene flow can also be estimated by dense interpolation of sparse matches. To this end, we find sparse matches across two stereo image pairs that are detected without any prior regularization and perform dense interpolation preserving geometric and motion boundaries by using edge information. A few iterations of variational energy minimization are performed to refine our results, which are thoroughly evaluated on the KITTI benchmark and additionally compared to state-of-the-art on MPI Sintel. For application in an automotive context, we further show that an optional ego-motion model helps to boost performance and blends smoothly into our approach to produce a segmentation of the scene into static and dynamic parts. René Schuster, Oliver Wasenmüller, Georg Kuschk, Christian Bailer, Didier Stricker |
WACV | 2 |
| 2017 | Time-of-flight sensor depth enhancement for automotive exhaust gasabstractThe Time-of-Flight (ToF) sensor has been envisioned as a candidate of next generation sensors for intelligent vehicles. One of the problems in automotive environment is that the sensor outputs wrong values if exhaust gas exists in the scene. In this paper, we provide two new contributions to the signal processing aspects of the ToF sensor for automotive use. First, we present the sensor characteristics and models for exhaust gas to cope with them. Second, we develop a depth enhancement algorithm to reject the influence of exhaust gas from multiple images. Experimental results demonstrate the effectiveness of the depth enhancement algorithm for both static data (including ground truth) and on-vehicle data acquired by the sensor mounted on a car. Tomonari Yoshida, Oliver Wasenmüller, Didier Stricker |
ICIP | 2 |
| 2016 | Augmented Reality 3D Discrepancy Check in Industrial ApplicationsabstractDiscrepancy check is a well-known task in industrial Augmented Reality (AR). In this paper we present a new approach consisting of three main contributions: First, we propose a new two-step depth mapping algorithm for RGB-D cameras, which fuses depth images with given camera pose in real-time into a consistent 3D model. In a rigorous evaluation with two public benchmarks we show that our mapping outperforms the state-of-the-art in accuracy. Second, we propose a semi-automatic alignment algorithm, which rapidly aligns a reference model to the reconstruction. Third, we propose an algorithm for 3D discrepancy check based on pre-computed distances. In a systematic evaluation we show the superior performance of our approach compared to state-of-the-art 3D discrepancy checks. Oliver Wasenmüller, Marcel Meyer, Didier Stricker |
ISMAR | 1 |
| 2016 | CoRBS: Comprehensive RGB-D benchmark for SLAM using Kinect v2abstractIn scientific evaluation public datasets and benchmarks are indispensable to perform objective assessment. In this paper we present a new Comprehensive RGB-D Benchmark for SLAM (CoRBS). In contrast to state-of-the-art RGB-D SLAM benchmarks, we provide the combination of real depth and color data together with a ground truth trajectory of the camera and a ground truth 3D model of the scene. Our novel benchmark allows for the first time to independently evaluate the localization as well as the mapping part of RGB-D SLAM systems with real data. We obtained the ground truth for the trajectory using an external motion capture system and for the scene geometry via an external 3D scanner, each with sub-millimeter precision. With precise calibration and systematic validation we ensured the high quality of CoRBS. Our dataset contains twenty image sequences of four different scenes captured with a Kinect v2. We provide all data in a global coordinate system to enable direct evaluation without any further alignment or calibration. Oliver Wasenmüller, Marcel Meyer, Didier Stricker |
WACV | 1 |