VLDB 2026 Research / reviewers in the wild / expert
Rudolf Mester
dblp:04/2727
· DBLP profile ↗
58ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0002-6932-0606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 8 since 2021Systems, architecture and hardware · 7 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Near-Shore Mapping for Detection and Tracking of VesselsabstractFor an autonomous surface vessel (ASV) to dock, it must track other vessels close to the docking area. Kayaks present a particular challenge due to their proximity to the dock and relatively small size. Maritime target tracking has typically employed land masking to filter out land and the dock. However, imprecise land masking makes it difficult to track close-to-dock objects. Our approach uses Light Detection And Ranging (LiDAR) data and maps the docking area before tracking. The precise 3D measurements allow for precise map creation. However, the mapping could result in static, yet potentially moving, objects being mapped. We detect and filter out potentially moving objects from the LiDAR data by utilizing image data. The visual vessel detection and segmentation method is a neural network that is trained on our labeled data. Close-to-shore tracking improves with an accurate map and is demonstrated on a recently gathered real-world dataset. The dataset contains multiple sequences of a kayak and a day cruiser moving close to the dock, in a collision path with an autonomous ferry prototype. Nicholas Dalhaug, Annette Stahl, Rudolf Mester, Edmund Førland Brekke |
FUSION | 3 |
| 2025 | Stixel-Based Free Space Estimation for USVs Using Stereo Camera and LiDARabstractUnmanned surface vehicles (USVs) require robust situational awareness to navigate safely in complex maritime environments. A critical element of this is to identify the free navigable space around the USV. Free water regions can be derived from water segmentation in the image. However, these segmented regions must be transformed into a bird's eye view (BEV) representation to be utilized effectively in motion planning. This paper proposes a novel approach to estimate free navigable space in a BEV format by integrating a stereo camera and light detection and ranging (LiDAR). The proposed method uses water segmentation to delineate the water surface and represents the closest obstacles in the USV line of sight using vertical planar rectangles known as Stixels. The depth of these Stixels is derived from LiDAR data, ensuring precise positioning in space. The effectiveness of the approach is demonstrated through experiments conducted on real-world data collected from the milliAmpere 2 (MA2) autonomous ferry prototype in Trondheim, Norway. Qualitative evaluations focusing on accuracy and temporal consistency confirm its ability to reliably detect free navigable areas in complex maritime environments. Johannes Robert Skarø, Trym Anthonsen Nygård, Rudolf Mester, Annette Stahl, Edmund Førland Brekke |
FUSION | 3 |
| 2025 | Visual Lidar Recursive Online Tracker (ViLiROT) for Autonomous Surface VesselsabstractWe propose a multi-sensor fusion pipeline for multiple object tracking in autonomous surface vessels using lidar and camera data. Our approach follows the tracking-by-detection paradigm, leveraging the precision of lidar for accurate state estimation and camera data for robust association. The method addresses issues with false tracks from lidar returns by suppressing non-moving objects on the basis of optical flow. We compare the proposed pipeline against prior work, particularly in the use of lidar and stereo cameras as depth modalities, demonstrating its effectiveness in improving tracking performance. Henrik Hilmarsen, Nicholas Dalhaug, Trym Anthonsen Nygård, Edmund Førland Brekke, Annette Stahl, Rudolf Mester |
ICRA | 6 |
| 2024 | Combining Short and Wide Baseline Stereo Cameras for Improved Maritime Target TrackingabstractTarget tracking is essential for autonomous vehicles to avoid collisions. Using a stereo camera for the target tracking gives a dense representation of the targets, contrary to the the sparser data on typical radars and lidars. With a wider baseline stereo camera the depth measurements are more accurate, but the stereo matching challenge is greater, especially in the maritime domain with reflections on the water. Earlier classical methods of tracking using stereo cameras have often tracked targets by first doing water surface estimation and then finding objects perturbing the plane. The challenge is then to get a good estimate of the water surface plane while still having precise measurements to the targets. We propose both a short baseline method and a multi-baseline method for target detection. The multi-baseline method uses a short baseline stereo camera to find the water plane and uses a wider baseline stereo camera to get accurate target measurements. The targets are consistently being tracked when using data collected during the summer of 2023 from an autonomous ferry prototype compared to ground truth GNSS tracks. The short baseline method achieves minimal error for a day cruiser boat 40 m away using a camera baseline of only 12 cm. The multi-baseline method further improves the accuracy of boat measurements, especially for a far-away small kayak. Nicholas Dalhaug, Annette Stahl, Rudolf Mester, Edmund Førland Brekke |
FUSION | 3 |
| 2024 | FusedWSS: Water Surface Segmentation Fusing Machine Learning and Geometric CuesabstractNavigating unmanned surface vehicles (USVs) in urban waterways presents unique challenges due to irregular waterlines, obstacles, and reflections in the water. Determining the collision-free navigable area is crucial to enable safe USV operation. This paper introduces Fused Water Surface Segmentation (FusedWSS), a novel approach to water surface segmentation that aims to enhance navigation capabilities for USVs in complex harbor environments using a stereo camera. The method locates the water plane by performing plane fitting with outlier rejection and plane validation on the reconstructed 3D point cloud. From the plane parameters, the virtual horizon line is inferred and used for point cloud and image cropping. The water surface mask and virtual horizon line are fused with a deep learningbased semantic segmentation method to produce accurate and reliable water masks for each image frame. Additional refinement of the water mask is performed using detected obstacle masks. Validation was carried out using data from the MilliAmpere 2 autonomous ferry prototype in Trondheim, Norway, and a publicly available maritime dataset, demonstrating the efficacy of the methods. Jon Torgeir Grini, Rudolf Mester, Trym Anthonsen Nygård, Nicholas Dalhaug, Edmund Førland Brekke, Annette Stahl |
FUSION | 2 |
| 2024 | Maritime Tracking-By-Detection with Object Mask Depth Retrieval Through Stereo Vision and LidarabstractThe momentum towards autonomous technology is building up in the maritime domain, as the automotive industry has made big steps towards autonomous driving. The automotive industry has increasingly utilized visual methods for multi-object tracking (MOT), with the help of accessible benchmarking datasets such as KITTI. This paper presents a tracking pipeline that tracks in the world frame by using elements of a well-established visual tracking method that tracks objects in the image frame. The pipeline fuses 3D information from lidar or stereo vision with object masks from a deep learning-based ship detector. To handle occlusions, we implemented a track manager that predicts lost objects’ movement until they reappear. Also, we provide a comparison between using lidar and stereo as the depth modality in the tracking pipeline. Results from a real-world experiment indicate that camera-lidar fusion gives consistently precise estimates, while the precision with stereo depends on the range and the type of vessel tracked. Henrik Hilmarsen, Nicholas Dalhaug, Trym Anthonsen Nygård, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 5 |
| 2024 | A General Low-Parameter 3D Ship Hull Extent Model for Object TrackingabstractIn autonomous vehicle systems, it is paramount to detect other objects in the vicinity and track their movement. Extended Object Tracking (EOT) provides a convenient framework for tracking objects using high-resolution sensor data by defining models for the object’s spatial dimensions (a.k.a. extent). In maritime applications, the objects of interest are mainly other maritime vessels, and these vary greatly in shape and size. This diversity proves to be a challenge for defining general extent models that both give accurate representations for most vessels and that do not depend on a large number of parameters. In this paper, a general three-dimensional low-parameter ship hull model designed for EOT is presented. The presented extent model is constructed by intertwining a polynomial representation along the vertical direction with a frequency representation along the horizontal plane. However, to reduce the dimension of the parameter space without compromising its accuracy, the horizontal frequency representation is modified by performing a Principal Component Analysis (PCA). In particular, this extent representation does not require an underlying discretization grid, which makes the model scalable and therefore well-suited for modeling objects that vary greatly in size. Michael Ernesto López, Kjetil Vasstein, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 4 |
| 2024 | RGB-D Mapping and Tracking in a Plenoxel Radiance FieldabstractThe widespread adoption of Neural Radiance Fields (NeRFs) have ensured significant advances in the domain of novel view synthesis in recent years. These models capture a volumetric radiance field of a scene, creating highly convincing, dense, photorealistic models through the use of simple, differentiable rendering equations. Despite their popularity, these algorithms suffer from severe ambiguities in visual data inherent to the RGB sensor, which means that although images generated with view synthesis can visually appear very believable, the underlying 3D model will often be wrong. This considerably limits the usefulness of these models in practical applications like Robotics and Extended Reality (XR), where an accurate dense 3D reconstruction otherwise would be of significant value. In this paper, we present the vital differences between view synthesis models and 3D reconstruction models. We also comment on why a depth sensor is essential for modeling accurate geometry in general outward-facing scenes using the current paradigm of novel view synthesis methods. Focusing on the structure-from-motion task, we practically demonstrate this need by extending the Plenoxel radiance field model: Presenting an analytical differential approach for dense mapping and tracking with radiance fields based on RGB-D data without a neural network. Our method achieves state-of-the-art results in both mapping and tracking tasks, while also being faster than competing neural network-based approaches. The code is available at: https://github.com/ysus33/RGB-D_Plenoxel_Mapping_Tracking.git. Andreas Langeland Teigen, Yeonsoo Park, Annette Stahl, Rudolf Mester |
WACV | 4 |
| 2023 | Maritime radar odometry inspired by visual odometryabstractFuture autonomous ships will need several redundant positioning systems to navigate reliably. Global Navigation Satellite Systems are highly accurate but they are susceptible to disruptions and intentional jamming. Maritime radars have long range and are robust against bad weather and darkness, but the use for ownship motion estimation has received relatively little attention in the research field. In this work, we present a radar odometry estimation method inspired by advances in visual odometry and simultaneous localization and mapping. The method works on raw radar data in a coastal environment and combines the Kanade-Lucas-Tomashi tracker with a factor graph back-end. We test it on data from a large ship with a maritime radar with a range of 19 km. We find that it is robust with only a small drift and no erroneous jumps in the estimate. Henrik D. Flemmen, Rudolf Mester, Annette Stahl, Torleiv H. Bryne, Edmund Førland Brekke |
FUSION | 2 |
| 2023 | Multiscan Shape Estimation for Extended Object TrackingabstractExtended Object Tracking (EOT) is a advantageous technique for achieving situational awareness in autonomous vehicle systems. The EOT problem is to both estimate the movement and spatial dimensions of an object using high-resolution measurements. In the case of laser measurements or other types of measurements that correspond to points on the object’s boundary, the true measurement model of the EOT problem is based on an implicit equation for the measurement coordinates. This intrinsic implicity is often not addressed directly in several EOT models found in the literature. In this paper, the EOT problem is reformulated as a least square minimization problem without compromising the original implicit measurement model by introducing an extra variable for each measurement. In addition, this new least squares formulation allows considering measurements and state variables for a whole time window, and not just a single time step. An EOT algorithm based on solving the derived least squares minimization problem is proposed and tested with simulated scenarios. Michael Ernesto López, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 3 |
| 2023 | Realistic Full-Body Anonymization with Surface-Guided GANsabstractRecent work on image anonymization has shown that generative adversarial networks (GANs) can generate near-photorealistic faces to anonymize individuals. However, scaling up these networks to the entire human body has remained a challenging and yet unsolved task. We propose a new anonymization method that generates realistic humans for in-the-wild images. A key part of our design is to guide adversarial nets by dense pixel-to-surface correspondences between an image and a canonical 3D surface. We introduce Variational Surface-Adaptive Modulation (V-SAM) that embeds surface information throughout the generator. Combining this with our novel discriminator surface supervision loss, the generator can synthesize high quality humans with diverse appearances in complex and varying scenes. We demonstrate that surface guidance significantly improves image quality and diversity of samples, yielding a highly practical generator. Finally, we show that our method preserves data usability without infringing privacy when collecting image datasets for training computer vision models. Source code and appendix is available at: github.com/hukkelas/full_body_anonymization Håkon Hukkelås, Morten Smebye, Rudolf Mester, Frank Lindseth |
WACV | 3 |
| 2022 | HD Ground - A Database for Ground Texture Based LocalizationabstractWe present the HD Ground Database, a comprehensive database for ground texture based localization. It contains sequences of a variety of textures, obtained using a downward facing camera. In contrast to existing databases of ground images, the HD Ground Database is larger, has a greater variety of textures, and has a higher image resolution with less motion blur. Also, our database enables the first systematic study of how natural changes of the ground that occur over time affect localization performance, and it allows to examine a teach-and-repeat navigation scenario. We use the HD Ground Database to evaluate four state-of-the-art localization approaches for global localization, localization with the approximate pose being known, and relative localization. Jan Fabian Schmid, Stephan F. Simon, Raaghav Radhakrishnan, Simone Frintrop, Rudolf Mester |
ICRA | 5 |
| 2022 | Geolocation estimation of target vehicles using image processing and geometric computationabstractEstimating vehicles’ locations is one of the key components in intelligent traffic management systems (ITMSs) for increasing traffic scene awareness. Traditionally, stationary sensors have been employed in this regard. The development of advanced sensing and communication technologies on modern vehicles (MVs) makes it feasible to use such vehicles as mobile sensors to estimate the traffic data of observed vehicles. This study aims to explore the capabilities of a monocular camera mounted on an MV in order to estimate the geolocation of the observed vehicle in a global positioning system (GPS) coordinate system. We proposed a new methodology by integrating deep learning, image processing, and geometric computation to address the observed-vehicle localization problem. To evaluate our proposed methodology, we developed new algorithms and tested them using real-world traffic data. The results indicated that our proposed methodology and algorithms could effectively estimate the observed vehicle’s latitude and longitude dynamically. Elnaz Namazi, Rudolf Mester, Chaoru Lu, Jingyue Li |
Neurocomputing | 2 |
| 2021 | Urban Traffic Surveillance (UTS): A fully probabilistic 3D tracking approach based on 2D detectionsabstractUrban Traffic Surveillance (UTS) is a surveillance system based on a monocular and calibrated video camera that detects vehicles in an urban traffic scenario with dense traffic on multiple lanes and vehicles performing sharp turning maneuvers. UTS then tracks the vehicles using a 3D bounding box representation and a physically reasonable 3D motion model relying on an unscented Kalman filter based approach. Since UTS recovers positions, shape and motion information in a three-dimensional world coordinate system, it can be employed to recognize diverse traffic violations or to supply intelligent vehicles with valuable traffic information. We build on YOLOv3 as a detector yielding 2D bounding boxes and class labels for each vehicle. A 2D detector renders our system much more independent to different camera perspectives as a variety of labeled training data is available. This allows for a good generalization while also being more hardware efficient. The task of 3D tracking based on 2D detections is supported by integrating class specific prior knowledge about the vehicle shape. We quantitatively evaluate UTS using self generated synthetic data and ground truth from the CARLA simulator, due to the non-existence of datasets with an urban vehicle surveillance setting and labeled 3D bounding boxes. Additionally, we give a qualitative impression of how UTS performs on real-world data. Our implementation is capable of operating in real time on a reasonably modern workstation. To the best of our knowledge, UTS is to date the only 3D vehicle tracking system in a surveillance scenario (static camera observing moving targets). Henry Bradler, Adrian Kretz, Rudolf Mester |
IV | 3 |
| 2020 | Ground Texture Based Localization Using Compact Binary DescriptorsabstractGround texture based localization is a promising approach to achieve high-accuracy positioning of vehicles. We present a self-contained method that can be used for global localization as well as for subsequent local localization updates, i.e. it allows a robot to localize without any knowledge of its current whereabouts, but it can also take advantage of a prior pose estimate to reduce computation time significantly. Our method is based on a novel matching strategy, which we call identity matching, that is based on compact binary feature descriptors. Identity matching treats pairs of features as matches only if their descriptors are identical. While other methods for global localization are faster to compute, our method reaches higher localization success rates, and can switch to local localization after the initial localization. Jan Fabian Schmid, Stephan F. Simon, Rudolf Mester |
ICRA | 3 |
| 2020 | Ground Texture Based Localization: Do We Need to Detect Keypoints?abstractLocalization using ground texture images recorded with a downward-facing camera is a promising approach to achieve reliable high-accuracy vehicle positioning. A common way to accomplish the task is to focus on prominent features of the ground texture such as stones and cracks. Our results indicate that with an approximately known camera pose it is sufficient to use arbitrary ground regions, i.e. extracting features at random positions without significant loss in localization performance. Additionally, we propose a real-time capable CPU-only localization method based on this idea, and suggest possible improvements for further research. Jan Fabian Schmid, Stephan F. Simon, Rudolf Mester |
IROS | 3 |
| 2019 | Features for Ground Texture Based Localization - A Survey
Jan Fabian Schmid, Stephan F. Simon, Rudolf Mester |
BMVC | 3 |
| 2019 | Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesabstractExisting 3D scene flow estimation methods provide the 3D geometry and 3D motion of a scene and gain a lot of interest, for example in the context of autonomous driving. These methods are traditionally based on a temporal series of stereo images. In this paper, we propose a novel monocular 3D scene flow estimation method, called Mono-SF. Mono-SF jointly estimates the 3D structure and motion of the scene by combining multi-view geometry and single-view depth information. Mono-SF considers that the scene flow should be consistent in terms of warping the reference image in the consecutive image based on the principles of multi-view geometry. For integrating single-view depth in a statistical manner, a convolutional neural network, called ProbDepthNet, is proposed. ProbDepthNet estimates pixel-wise depth distributions from a single image rather than single depth values. Additionally, as part of ProbDepthNet, a novel recalibration technique for regression problems is proposed to ensure well-calibrated distributions. Our experiments show that Mono-SF outperforms state-of-the-art monocular baselines and ablation studies support the Mono-SF approach and ProbDepthNet design. Fabian Brickwedde, Steffen Abraham, Rudolf Mester |
ICCV | 3 |
| 2018 | Mono-Stixels: Monocular Depth Reconstruction of Dynamic Street ScenesabstractIn this paper we present mono-stixels, a compact environment representation specially designed for dynamic street scenes. Mono-stixels are a novel approach to estimate stixels from a monocular camera sequence instead of the traditionally used stereo depth measurements. Our approach jointly infers the depth, motion and semantic information of the dynamic scene as a 1D energy minimization problem based on optical flow estimates, pixel-wise semantic segmentation and camera motion. The optical flow of a stixel is described by a homography. By applying the mono-stixel model the degrees of freedom of a stixel-homography are reduced to only up to two degrees of freedom. Furthermore, we exploit a scene model and semantic information to handle moving objects. In our experiments we use the public available DeepFlow for optical flow estimation and FCN8s for the semantic information as inputs and show on the KITTI 2015 dataset that mono-stixels provide a compact and reliable depth reconstruction of both the static and moving parts of the scene. Thereby, mono-stixels overcome the limitation to static scenes of previous structure-from-motion approaches. Fabian Brickwedde, Steffen Abraham, Rudolf Mester |
ICRA | 3 |
| 2018 | CNN-based multi-frame IMO detection from a monocular cameraabstractThis paper presents a method for detecting independently moving objects (IMOs) from a monocular camera mounted on a moving car. A CNN-based classifier is employed to generate IMO candidate patches; independent motion is detected by geometric criteria on keypoint trajectories in these patches. Instead of looking only at two consecutive frames, we analyze keypoints inside the IMO candidate patches through multi-frame epipolar consistency checks. The obtained motion labels (IMO/static) are then propagated over time using the combination of motion cues and appearance-based information of the IMO candidate patches. We evaluate the performance of our method on the KITTI dataset, focusing on sub-sequences containing IMOs. Nolang Fanani, Matthias Ochs, Alina Sturck, Rudolf Mester |
Intelligent Vehicles Symposium | 4 |
| 2018 | Spatio-Temporal Depth Interpolation (STDI)abstractIn the area of autonomous driving, sensing the environment is most important for self-localization and egomotion estimation. Visual odometry/SLAM methods have proven capable to achieve good results, even in real-time applications by operating in a sparse mode. Running on a sequence, these methods need to continuously incorporate new features well distributed over the image. Therefore, the performance of these methods can be further improved, if they are supplied with coarse but dense initial depth information, that can be utilized at arbitrary sparse image positions. Previously triangulated depths and even high quality depth measurements of a LIDAR sensor are not suitable for this task, since they only provide a sparse depth map. To solve this issue, we introduce a novel interpolation method called Spatio-Temporal Depth Interpolation (STDI), which exploits spatial and temporal correlations of the data (e.g. sequences of sparse depth maps) to give a consistent dense output including associated uncertainties. STDI is a fused approach, which makes use of the most important components of a principal component analysis (PCA) (spatial information) and additionally is capable to re-use information of previously interpolated depth maps in a regression based approach (temporal information). We evaluate the quality of STDI on the KITTI visual odometry benchmark, where a sequence of extremely sparsely sampled depth maps (≈ 40 depth values) is densified and on the KITTI depth completion benchmark. The latter deals with the densification of sparse LIDAR input. Of course, our method is not limited to these applications and can be used for any densification of sparse sequential data which is expected to contain spatial and/or temporal correlations (e.g. initialization for dense optical flow methods based on a sparse measurement). Matthias Ochs, Henry Bradler, Rudolf Mester |
Intelligent Vehicles Symposium | 3 |
| 2017 | Multimodal scale estimation for monocular visual odometryabstractMonocular visual odometry / SLAM requires the ability to deal with the scale ambiguity problem, or equivalently to transform the estimated unscaled poses into correctly scaled poses. While propagating the scale from frame to frame is possible, it is very prone to the scale drift effect. We address the problem of monocular scale estimation by proposing a multimodal mechanism of prediction, classification, and correction. Our scale correction scheme combines cues from both dense and sparse ground plane estimation; this makes the proposed method robust towards varying availability and distribution of trackable ground structure. Instead of optimizing the parameters of the ground plane related homography, we parametrize and optimize the underlying motion parameters directly. Furthermore, we employ classifiers to detect scale outliers based on various features (e.g. moments on residuals). We test our method on the challenging KITTI dataset and show that the proposed method is capable to provide scale estimates that are on par with current state-of-the-art monocular methods without using bundle adjustment or RANSAC. Nolang Fanani, Alina Sturck, Marc Barnada, Rudolf Mester |
Intelligent Vehicles Symposium | 4 |
| 2017 | Learning rank reduced interpolation with principal component analysisabstractMost iterative optimization algorithms for motion, depth estimation or scene reconstruction, both sparse and dense, rely on a coarse and reliable dense initialization to bootstrap their optimization procedure. This makes techniques important that allow to obtain a dense but still approximative representation of a desired 2D structure (e.g., depth maps, optical flow, disparity maps) from a very sparse measurement of this structure. The method presented here exploits the complete information given by the principal component analysis (PCA), the principal basis and its prior distribution. The method is able to determine a dense reconstruction even if only a very sparse measurement is available. When facing such situations, typically the number of principal components is further reduced which results in a loss of expressiveness of the basis. We overcome this problem and inject prior knowledge in a maximum a posteriori (MAP) approach. We test our approach on the KITTI and the Virtual KITTI dataset and focus on the interpolation of depth maps for driving scenes. The evaluation of the results shows good agreement to the ground truth and is clearly superior to the results of an interpolation by the nearest neighbor method which disregards statistical information. Matthias Ochs, Henry Bradler, Rudolf Mester |
Intelligent Vehicles Symposium | 3 |
| 2017 | Joint Epipolar Tracking (JET): Simultaneous Optimization of Epipolar Geometry and Feature CorrespondencesabstractTraditionally, pose estimation is considered as a two step problem. First, feature correspondences are determined by direct comparison of image patches, or by associating feature descriptors. In a second step, the relative pose and the coordinates of corresponding points are estimated, most often by minimizing the reprojection error (RPE). RPE optimization is based on a loss function that is merely aware of the feature pixel positions but not of the underlying image intensities. In this paper, we propose a sparse direct method which introduces a loss function that allows to simultaneously optimize the unscaled relative pose, as well as the set of feature correspondences directly considering the image intensity values. Furthermore, we show how to integrate statistical prior information on the motion into the optimization process. This constructive inclusion of a Bayesian bias term is particularly efficient in application cases with a strongly predictable (short term) dynamic, e.g. in a driving scenario. In our experiments, we demonstrate that the 'JET' algorithm we propose outperforms the classical reprojection error optimization on two synthetic datasets and on the KITTI dataset. The JET algorithm runs in real-time on a single CPU thread. Henry Bradler, Matthias Ochs, Nolang Fanani, Rudolf Mester |
WACV | 4 |
| 2017 | Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment?
Nolang Fanani, Alina Sturck, Matthias Ochs, Henry Bradler, Rudolf Mester |
Image Vis. Comput. | 5 |
| 2016 | Lost and Found: detecting small road hazards for self-driving vehiclesabstractDetecting small obstacles on the road ahead is a critical part of the driving task which has to be mastered by fully autonomous cars. In this paper, we present a method based on stereo vision to reliably detect such obstacles from a moving vehicle. Peter Pinggera, Sebastian Ramos, Stefan K. Gehrig, Uwe Franke, Carsten Rother, Rudolf Mester |
IROS | 6 |
| 2016 | Evaluating visual ADAS components on the COnGRATS datasetabstractWe present a framework that supports the development and evaluation of vision algorithms in the context of driver assistance applications and traffic surveillance. This framework allows the creation of highly realistic image sequences featuring traffic scenarios. The sequences are created with a realistic state of the art vehicle physics model; different kinds of environments are featured, thus providing a wide range of testing scenarios. Due to the physically-based rendering technique and variable camera models employed for the image rendering process, we can simulate different sensor setups and provide appropriate and fully accurate ground truth data. Daniel H. Biedermann, Matthias Ochs, Rudolf Mester |
Intelligent Vehicles Symposium | 3 |
| 2016 | Keypoint trajectory estimation using propagation based trackingabstractOne of the major steps in visual environment perception for automotive applications is to track keypoints and to subsequently estimate egomotion and environment structure from the trajectories of these keypoints. This paper presents a propagation based tracking method to obtain the 2D trajectories of keypoints from a sequence of images in a monocular camera setup. Instead of relying on the classical RANSAC to obtain accurate keypoint correspondences, we steer the search for keypoint matches by means of propagating the estimated 3D position of the keypoint into the next frame and verifying the photometric consistency. In this process, we continuously predict, estimate and refine the frame-to-frame relative pose which induces the epipolar relation. Experiments on the KITTI dataset as well as on the synthetic COnGRATS dataset show promising results on the estimated courses and accurate keypoint trajectories. Nolang Fanani, Matthias Ochs, Henry Bradler, Rudolf Mester |
Intelligent Vehicles Symposium | 4 |
| 2015 | Learning relative photometric differences of pairs of camerasabstractWe present an approach to learn relative photometric differences between pairs of cameras, which have partially overlapping fields of views. This is an important problem, especially in appearance based methods to correspondence estimation or object identification in multi-camera systems where grey values observed by different cameras are processed. We model intensity differences among pairs of cameras by means of a low order polynomial (Gray Value Transfer Function - GVTF ) which represents the characteristic curve of the mapping of grey values siproduced by camera Cito the corresponding grey values sjacquired with camera Cj. While the estimation of the GVTF parameters is straight forward once a set of truly corresponding pairs of grey values is available, the non trivial task in the GVTF estimation process solved in this paper is the extraction of corresponding grey value pairs in the presence of geometric and photometric errors. We also present a temporal GVTF update scheme to adapt to gradual global illumination changes, e.g., due to the change of daylight. Christian Conrad, Rudolf Mester |
AVSS | 2 |
| 2015 | High-performance long range obstacle detection using stereo visionabstractReliable detection of obstacles at long range is crucial for the timely response to hazards by fast-moving safety-critical platforms like autonomous cars. We present a novel method for the joint detection and localization of distant obstacles using a stereo vision system on a moving platform. The approach is applicable to both static and moving obstacles and pushes the limits of detection performance as well as localization accuracy. Peter Pinggera, Uwe Franke, Rudolf Mester |
IROS | 3 |
| 2015 | Estimation of automotive pitch, yaw, and roll using enhanced phase correlation on multiple far-field windowsabstractThe online-estimation of yaw, pitch, and roll of a moving vehicle is an important ingredient for systems which estimate egomotion, and 3D structure of the environment in a moving vehicle from video information. We present an approach to estimate these angular changes from monocular visual data, based on the fact that the motion of far distant points is not dependent on translation, but only on the current rotation of the camera. The presented approach does not require features (corners, edges, …) to be extracted. It allows to estimate in parallel also the illumination changes from frame to frame, and thus allows to largely stabilize the estimation of image correspondences and motion vectors, which are most often central entities needed for computating scene structure, distances, etc. The method is significantly less complex and much faster than a full egomotion computation from features, such as PTAM [6], but it can be used for providing motion priors and reduce search spaces for more complex methods which perform a complete analysis of egomotion and dynamic 3D structure of the scene in which a vehicle moves. Marc Barnada, Christian Conrad, Henry Bradler, Matthias Ochs, Rudolf Mester |
Intelligent Vehicles Symposium | 5 |
| 2015 | Robust stereo visual odometry from monocular techniquesabstractVisual odometry is one of the most active topics in computer vision. The automotive industry is particularly interested in this field due to the appeal of achieving a high degree of accuracy with inexpensive sensors such as cameras. The best results on this task are currently achieved by systems based on a calibrated stereo camera rig, whereas monocular systems are generally lagging behind in terms of performance. We hypothesise that this is due to stereo visual odometry being an inherently easier problem, rather than than due to higher quality of the state of the art stereo based algorithms. Under this hypothesis, techniques developed for monocular visual odometry systems would be, in general, more refined and robust since they have to deal with an intrinsically more difficult problem. In this work we present a novel stereo visual odometry system for automotive applications based on advanced monocular techniques. We show that the generalization of these techniques to the stereo case result in a significant improvement of the robustness and accuracy of stereo based visual odometry. We support our claims by the system results on the well known KITTI benchmark, achieving the top rank for visual only systems*. Mikael Persson, Tommaso Piccini, Michael Felsberg, Rudolf Mester |
Intelligent Vehicles Symposium | 4 |
| 2015 | Enhanced Phase Correlation for Reliable and Robust Estimation of Multiple Motion Distributions
Matthias Ochs, Henry Bradler, Rudolf Mester |
PSIVT | 3 |
| 2014 | Know Your Limits: Accuracy of Long Range Stereoscopic Object Measurements in Practice
Peter Pinggera, David Pfeiffer, Uwe Franke, Rudolf Mester |
ECCV (2) | 4 |
| 2014 | When Patches Match - A Statistical View on Matching under Illumination VariationabstractWe discuss matching measures (scores and residuals) for comparing image patches under unknown affine photometric (=intensity) transformations. In contrast to existing methods, we derive a fully symmetric matching measure which reflects the fact that both copies of the signal are affected by measurement errors ('noise'), not only one. As it turns out, this evolves into an Eigen system problem, however a simple direct solution for all entities of interest can be given. We strongly advocate for constraining the estimated gain ratio and the estimated mean value offset to realistic ranges, thus preventing the matching scheme from locking into unrealistic correspondences. Rudolf Mester, Christian Conrad |
ICPR | 1 |
| 2014 | Spider-based Stixel object segmentationabstractStereo vision has established in the field of driver assistance and vehicular safety systems. Next steps along the road towards accident free driving aim to assist the driver in increasingly complex situations such as inner-city traffic. In order to achieve these goals, it is desirable to incorporate higher-order object knowledge in the stereo vision-based understanding of traffic scenes. In particular, object shape and dimension information can help to achieve correct interpretations. However, typically this kind of higher-order information results in a difficult energy minimization problem since large areas of the input image have to be constrained. In this contribution, an efficient global optimization approach based on dynamic programming is proposed that is able to take into account such higher-order object knowledge. The approach is built upon a simple tree representation of the Dynamic Stixel World, an efficient super-pixel object representation. Experiments show that object segmentation can be improved significantly by means of the higher-order object information. Friedrich Erbs, Andreas Witte, Timo Scharwächter, Rudolf Mester, Uwe Franke |
Intelligent Vehicles Symposium | 4 |
| 2013 | Reducing camera vibrations and photometric changes in surveillance videoabstractWe analyze the consequences of instabilities and fluctuations, such as camera shaking and illumination/exposure changes, on typical surveillance video material and devise a systematic way to compensate these changes as much as possible. The phase correlation method plays a decisive role in the proposed scheme, since it is inherently insensitive to gain and offset changes, as well as insensitive against different linear degradations (due to time-variant motion blur) in subsequent images. We show that the listed variations can be compensated effectively, and the image data can be equilibrated significantly before a temporal change detection and/or a background-based detection is performed. We verify the usefulness of the method by comparative tests with and without stabilization, using the changedetection.net benchmark and several state-of-the-art detections methods. Jens Eisenbach, Matthias Mertz, Christian Conrad, Rudolf Mester |
AVSS | 4 |
| 2013 | Discriminative Subspace ClusteringabstractWe present a novel method for clustering data drawn from a union of arbitrary dimensional subspaces, called Discriminative Subspace Clustering (DiSC). DiSC solves the subspace clustering problem by using a quadratic classifier trained from unlabeled data (clustering by classification). We generate labels by exploiting the locality of points from the same subspace and a basic affinity criterion. A number of classifiers are then diversely trained from different partitions of the data, and their results are combined together in an ensemble, in order to obtain the final clustering result. We have tested our method with 4 challenging datasets and compared against 8 state-of-the-art methods from literature. Our results show that DiSC is a very strong performer in both accuracy and robustness, and also of low computational complexity. Vasileios Zografos, Liam F. Ellis, Rudolf Mester |
CVPR | 3 |
| 2012 | Learning geometrical transforms between multi camera views using Canonical Correlation AnalysisabstractWe present an unsupervised and sampling-free approach to learn the correspondence relations between pairs of cameras in closed form employing a linear model known as Canonical Correlation Analysis (CCA). The only assumption we make is that the relative orientation between the cameras involved is fixed. In a two stage algorithm, we first learn the inter-image transformation based on CCA. This analysis usually has to be done in a multi-scale framework, as applying CCA directly to full resolution images may be computationally prohibitive. In the second stage we employ the learnt transformation which is given only implicitly and predict for a given pixel in a first view its corresponding region within a second view. We denote these regions as correspondence prior. CCA has been introduced by Hotelling [2] as a method of analyzing the relations between two sets of variates and can be applied in closed form. Consider two random vectors x and y where x 2 R N and y 2 R M . The goal of CCA is to find basis vectors for which the correlation between x and y when projected onto the basis vectors are mutually maximized [3]. In the case of a single pair of basis vectors u 2 R N ,v 2 R M the projections are given as a=u T x and b=v T y. Assuming E[x]=E[y]=0, the correlation r between a and b can be written as: Christian Conrad, Rudolf Mester |
BMVC | 2 |
| 2012 | Illumination invariance for driving scene optical flow using comparagram preselectionabstractIn the recent years, advanced video sensors have become common in driver assistance, coping with the highly dynamic lighting conditions by nonlinear exposure adjustments. However, many computer vision algorithms are still highly sensitive to the resulting sudden brightness changes. We present a method that is able to estimate the relative intensity transfer function (RITF) between images in a sequence even for moving cameras. The according compensation of the input images can improve the performance of further vision tasks significantly, here demonstrated by results from optical flow. Our method identifies corresponding intensity values from areas in the images where no apparent motion is present. The RITF is then estimated from that data and regularized based on its curvature. Finally, built-in tests reliably flag image pairs with `adverse conditions' where no compensation could be performed. David Dederscheck, Thomas Müller 0008, Rudolf Mester |
Intelligent Vehicles Symposium | 3 |
| 2011 | Boosting segmentation results by contour relaxationabstractThis paper presents a versatile algorithmic building block that allows to significantly improve intermediate and final results of numerous variations of segmentation. The segmentation `context' can be very different in terms of the used data modality (gray scale, color, texture features, depth data, motion, ...), in terms of single frame vs. sequence segmentation, and in terms of the used initialization (measurement space clustering vs. `blind' initialization vs. interactively `sketching' the segmentation). For all these mentioned variations, the contour relaxation approach presented here offers the capability of very efficiently obtaining a segmentation result that is both visually pleasing as well as locally optimal with respect to a statistically well justified target functional. Alvaro Guevara, Christian Conrad, Rudolf Mester |
ICIP | 3 |
| 2010 | Optical Rails: View-Based Track Following with Hemispherical Environment Model and Orientation View DescriptorsabstractWe present a purely view-based method for robot navigation along a prerecorded track using compact omni directional view-descriptors. This paper focuses on a new model for the navigation environment to determine the steering direction by efficient holistic comparison of views. The concept of view descriptors based on low-order expansion of local orientation vectors into spherical harmonic basis functions is augmented by a linear illumination model, providing discriminative view matching also under illumination changes. David Dederscheck, Martin Zahn, Holger Friedrich 0001, Rudolf Mester |
ICPR | 4 |
| 2008 | A Statistical Confidence Measure for Optical Flows
Claudia Kondermann, Rudolf Mester, Christoph S. Garbe |
ECCV (3) | 2 |
| 2007 | Highly Accurate Orientation Estimation using Steerable FiltersabstractThis paper presents a generalized theoretical framework for designing accurate steerable filters for orientation estimation. We derive the necessary properties of orientation estimation filters in their most general form. Based on our framework, we implemented a highly angular-specific filter. Numerical experiments show the enhanced accuracy of our proposed filter as compared with existing filters. Jakub K. Kominiarczuk, Kai Krajsek, Rudolf Mester |
ICIP (5) | 3 |
| 2006 | A Maximum Likelihood Estimator for Choosing the Regularization Parameters in Global Optical Flow MethodsabstractGlobal optical flow estimation methods based on variational calculus contain a regularization parameter which controls the tradeoff between the different constraints on the optical flow field. The counterpart to the regularization parameter are the hyper-parameters in the Bayesian framework. These hyper-parameters have distinct physical meanings and thus can be inferred from the observable data. We derive a combined marginal maximum likelihood/maximum a posteriori (MML/MAP) estimator for simultaneously estimating hyper-parameters and optical flow for all differential variational approaches directly from the observed signal without any prior knowledge of the optical flow. Experiments demonstrate the performance of this optimization technique and show that the choice of the regularization parameter is an essential key-point in order to obtain precise motion estimation. Kai Krajsek, Rudolf Mester |
ICIP | 2 |
| 2003 | The generalization, optimization, and information-theoretic justification of filter-based and autocovariance-based motion estimationabstractWe discuss the theoretical foundations of measuring motion in video data, and relate this strongly to statistical estimation theory. A very general class of motion estimation methods is characterized by determining second order moments of filter bank outputs. These moments are represented in tensors, and motion estimation boils down to analyzing their eigensystems. An alternative approach is to directly estimate and analyze the autocorrelation of the given signal. We provide motivation for developing these approaches further towards directional entropy rate criteria rather than rely on conventional directional smoothness criteria. This pa- per emphasizes that prior knowledge on the video signal (e.g. spatial autocovariance, distribution of expected motion speed, noise spectrum,...) should be integrated into the motion estimation procedure. Relations between different classes of motion algorithms (differential, tensor-based, steerable filters...) are discussed and perspectives for a unification and enhancement of such procedures are presented. Rudolf Mester |
ICIP (3) | 1 |
| 2001 | Bayesian illumination-invariant motion detectionabstractWe describe an algorithm for change detection which is insensitive to both slow and fast temporal variations of scene illumination. Our algorithm is based on statistical decision theory by using a Bayesian approach. The goal is to detect only temporal changes which are induced by true scene changes, like motion, but not changes due to varying illumination or noise. To this end, our algorithm uses a simple illumination model which is invariant to common camera nonlinearities like gamma-nonlinearity. This is combined with a model for the influence of noise as well as an a priori model for the expected properties of the sought change masks. Key ingredients of the resulting algorithm are a suitable test statistic and an adaptive threshold mechanism. As the algorithm can be applied in a noniterative manner, it is also computationally attractive. Til Aach, Lutz Dümbgen, Rudolf Mester, Daniel Toth 0002 |
ICIP (3) | 3 |
| 2001 | Improving motion and orientation estimation using an equilibrated total least squares approachabstractThis paper outlines the ubiquitous presence of generalized orientation (or subspace) estimation problems in image analysis. We show the potential sources of bias in naive approaches to directional estimation problems, discuss countermeasures against this bias, and point out the direct relation to the total least squares problem. An improved method (using TLS and "equilibration") for a precise direct motion estimation of planar objects (8 parameter motion model, homography estimation) concludes this paper. Matthias Mühlich, Rudolf Mester |
ICIP (2) | 2 |
| 2001 | A considerable improvement in non-iterative homography estimation using TLS and equilibration
Matthias Mühlich, Rudolf Mester |
Pattern Recognit. Lett. | 2 |
| 1999 | Estimating Consistent Motion from Three Views: An Alternative to Trifocal Analysis
Stefan Trautwein, Matthias Mühlich, Dirk Feiden, Rudolf Mester |
CAIP | 4 |
| 1998 | The Role of Total Least Squares in Motion Analysis
Matthias Mühlich, Rudolf Mester |
ECCV (2) | 2 |
| 1996 | Detection of moving objects using a robust displacement estimation including a statistical error analysisabstractA new technique for the detection and description of moving objects in natural scenes is presented which is based on an object-oriented, statistical multifeature analysis of video sequences. To cope with the problem that image signal changes can have causes other than object motion, additional features, viz, texture and motion beyond temporal signal differences, are extracted and evaluated in an object-oriented fashion. The adaptation of this method to normal fluctuations of the observed scene is performed by a time-recursive space-variant estimation of the temporal probability distributions of the features. Feature data which differ significantly from the estimated distributions are interpreted to be caused by moving objects. For motion feature extraction, a robust displacement estimation algorithm is applied which is oriented towards the joint estimation of displacement vectors and their corresponding reliability measures. The reliability measures judge object motion and control the alarm setting. The advantages of the presented object detection algorithm compared to mere change detection techniques are demonstrated by some experiments. Due to its capability to automatically learn the observed scene the calibration effort of the sensor is extremely small. Michael Hötter, Rudolf Mester |
ICPR | 2 |
| 1996 | Detection and description of moving objects by stochastic modelling and analysis of complex scenes
Michael Hötter, Rudolf Mester, Frank Müller 0002 |
Signal Process. Image Commun. | 2 |
| 1995 | On texture analysis: Local energy transforms versus quadrature filters
Til Aach, André Kaup, Rudolf Mester |
Signal Process. | 3 |
| 1993 | Statistical model-based change detection in moving video
Til Aach, André Kaup, Rudolf Mester |
Signal Process. | 3 |
| 1992 | Spectral Entropy-Activity Classification in Adaptive Transform CodingabstractA two-dimensional feature space for block classification which discriminates reliably between blocks that require very different processing is proposed. In combination with a block 'activity' measurement, the introduced 'spectral entropy' feature offers the possibility to stabilize the reconstruction quality of transform coding systems for each processed block on a high level. The classification is valid and useful for threshold-based and zonal coding schemes.> Rudolf Mester, Uwe Franke |
IEEE J. Sel. Areas Commun. | 1 |
| 1990 | Combined displacement estimation and segmentation of stereo image pairs based on Gibbs random fieldsabstractA technique for displacement estimation in stereo image pairs, which is closely coupled with a segmentation of the images into homogeneously displaced regions, is described. An arbitrary initial displacement field is modified by a displacement vector relaxation, evaluating a statistical optimization criterion. Since the displacements are considered vector by vector during the relaxation, the method provides accurate locations of the boundaries between differently displaced regions. The proposed technique derives its power from the fact that the spatial array of displacements-and hence the resulting segmentation-is modeled by a Gibbs-Markov random field.> Til Aach, André Kaup, Rudolf Mester |
ICASSP | 3 |
| 1989 | Top-down image segmentation using object detection and contour relaxationabstractA novel segmentation technique that starts with the whole image being a single region is presented. First, an object detection scheme, which marks those locations where local statistics deviate significantly from the overall statistics, provides location and approximate shapes of the major objects (regions) in the scene. Exact boundaries are subsequently obtained by a contour relaxation algorithm, which includes a general model for typical region shapes. Object detection and contour relaxation are repeated recursively until a stable segmentation result is achieved. Segmentation results are presented.> Til Aach, Uwe Franke, Rudolf Mester |
ICASSP | 3 |