VLDB 2026 Research / reviewers in the wild / expert
Sebastian Bullinger
dblp:197/9724
· DBLP profile ↗
11ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-1584-5319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Strike the Balance: On-the-Fly Uncertainty Based User Interactions for Long-Term Video Object Segmentation
Stéphane Vujasinovic, Stefan Becker, Sebastian Bullinger, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen |
ACCV (2) | 3 |
| 2024 | Statewide Visual Geolocalization in the Wild
Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen |
ECCV (36) | 2 |
| 2023 | READMem: Robust Embedding Association for a Diverse Memory in Unconstrained Video Object Segmentation
Stéphane Vujasinovic, Sebastian Bullinger, Stefan Becker, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen |
BMVC | 2 |
| 2023 | Uncertainty-Aware Vision-Based Metric Cross-View GeolocalizationabstractThis paper proposes a novel method for vision-based metric cross-view geolocalization (CVGL) that matches the camera images captured from a ground-based vehicle with an aerial image to determine the vehicle's geo-pose. Since aerial images are globally available at low cost, they represent a potential compromise between two established paradigms of autonomous driving, i.e. using expensive high-definition prior maps or relying entirely on the sensor data captured at runtime. We present an end-to-end differentiable model that uses the ground and aerial images to predict a probability distribution over possible vehicle poses. We combine multiple vehicle datasets with aerial images from orthophoto providers on which we demonstrate the feasibility of our method. Since the ground truth poses are often inaccurate w.r.t. the aerial images, we implement a pseudo-label approach to produce more accurate ground truth poses and make them publicly available. While previous works require training data from the target region to achieve reasonable localization accuracy (i.e. same-area evaluation), our approach overcomes this limitation and outperforms previous results even in the strictly more challenging cross-area case. We improve the previous state-of-the-art by a large margin even without ground or aerial data from the test region, which highlights the model's potential for global-scale application. We further integrate the uncertainty-aware predictions in a tracking framework to determine the vehicle's trajectory over time resulting in a mean position error on KITTI-360 of 0.78m. Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen |
CVPR | 2 |
| 2022 | Revisiting Click-Based Interactive Video Object SegmentationabstractWhile current methods for interactive Video Object Segmentation (iVOS) rely on scribble-based interactions to generate precise object masks, we propose a Click-based interactive Video Object Segmentation (CiVOS) framework to simplify the required user workload as much as possible. CiVOS builds on de-coupled modules reflecting user interaction and mask propagation. The interaction module converts click-based interactions into an object mask, which is then inferred to the remaining frames by the propagation module. Additional user interactions allow for a refinement of the object mask. The approach is extensively evaluated on the popular interactive DAVIS dataset, but with an inevitable adaptation of scribble-based interactions with click-based counterparts. We consider several strategies for generating clicks during our evaluation to reflect various user inputs and adjust the DAVIS performance metric to perform a hardware-independent comparison. The presented CiVOS pipeline achieves competitive results, although requiring a lower user workload. Stéphane Vujasinovic, Sebastian Bullinger, Stefan Becker, Norbert Scherer-Negenborn, Michael Arens, Rainer Stiefelhagen |
ICIP | 2 |
| 2022 | Continuous Self-Localization on Aerial Images Using Visual and Lidar SensorsabstractThis paper proposes a novel method for geo-tracking, i.e. continuous metric self-localization in outdoor environments by registering a vehicle's sensor information with aerial imagery of an unseen target region. Geo- tracking methods offer the potential to supplant noisy signals from global navigation satellite systems (GNSS) and expensive and hard to maintain prior maps that are typically used for this purpose. The proposed geo-tracking method aligns data from on-board cameras and lidar sensors with geo-registered orthophotos to continuously localize a vehicle. We train a model in a metric learning setting to extract visual features from ground and aerial images. The ground features are projected into a top-down perspective via the lidar points and are matched with the aerial features to determine the relative pose between vehicle and orthophoto. Our method is the first to utilize on-board cameras in an end-to-end differentiable model for metric self-localization on unseen orthophotos. It exhibits strong generalization, is robust to changes in the environment and requires only geo-poses as ground truth. We evaluate our approach on the KITTI-360 dataset and achieve a mean absolute position error (APE) of 0.94m. We further compare with previous approaches on the KITTI odometry dataset and achieve state-of-the-art results on the geo-tracking task.3 Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen |
IROS | 2 |
| 2019 | 3D Object Trajectory Reconstruction using Instance-Aware Multibody Structure from Motion and Stereo Sequence ConstraintsabstractThree-dimensional environment perception is a key element of autonomous driving and driver assistance systems. A common image based approach to determine three-dimensional scene information is stereo matching, which is limited by the stereo camera baseline. In contrast to stereo matching based methods, we present an approach to reconstruct three-dimensional object trajectories combining temporal adjacent views for object point triangulation. We track two-dimensional object shapes on pixel level exploiting instance-aware semantic segmentation techniques and optical flow cues. We apply Structure from Motion (SfM) to object and background images to determine initial camera poses relative to object instances as well as background structures and refine the initial SfM results by integrating stereo camera constraints using factor graphs. We compute object trajectories using stereo sequence constraints of object and background reconstructions. We show qualitative results using publicly available video data of driving sequences. Due to the lack of suitable ground truth, we create a synthetic benchmark dataset of stereo sequences with vehicles in urban environments. Our algorithm achieves an average trajectory error of 0.09 meter using the dataset. The dataset is on our website1publicly available. Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen |
IV | 1 |
| 2018 | Multispectral Matching using Conditional Generative Appearance ModelingabstractThe precise determination of correspondences between pairs of images is still a fundamental building block of many computer vision systems. Despite the maturity of modern feature matchers, multispectral methods are still lacking robustness and speed. We focus on the problem of finding point correspondences in a multispectral imaging setup. Most methods aim at invariant feature transforms (e.g. multi-modal descriptors) which come at the cost of reduced discriminance.We model the appearance change by learning an image transformation, which maps one image modality to the respective target image, conditioned on the data of the original spectral band. This approach is coupled with a pipeline of state of the art matching methods with view synthesis of increasing complexity and algorithm run-time.We evaluate the approach on a wide spectrum of multispectral datasets including near-infrared, color-infrared and night and day thermal infrared imagery. The proposed approach provides significant improvements in terms of speed and robustness compared to standard multi-modal registration approaches. In addition, the approach fits very well into existing system approaches by design.Applications are numerous and include multispectral sensor fusion, multispectral odometry systems, multispectral segmentation or multispectral super-resolution methods. Christoph Bodensteiner, Sebastian Bullinger, Michael Arens |
AVSS | 2 |
| 2018 | 3D Vehicle Trajectory Reconstruction in Monocular Video Data Using Environment Structure Constraints
Sebastian Bullinger, Christoph Bodensteiner, Michael Arens, Rainer Stiefelhagen |
ECCV (10) | 1 |
| 2017 | Instance flow based online multiple object trackingabstractWe present a method to perform online Multiple Object Tracking (MOT) of known object categories in monocular video data. Current Tracking-by-Detection MOT approaches build on top of 2D bounding box detections. In contrast, we exploit state-of-the-art instance aware semantic segmentation techniques to compute 2D shape representations of target objects in each frame. We predict position and shape of segmented instances in subsequent frames by exploiting optical flow cues. We define an affinity matrix between instances of subsequent frames which reflects locality and visual similarity. The instance association is solved by applying the Hungarian method. We evaluate different configurations of our algorithm using the MOT 2D 2015 train dataset. The evaluation shows that our tracking approach is able to track objects with high relative motions. In addition, we provide results of our approach on the MOT 2D 2015 test set for comparison with previous works. We achieve a MOTA score of 32.1. Sebastian Bullinger, Christoph Bodensteiner, Michael Arens |
ICIP | 1 |
| 2016 | Moving object reconstruction in monocular video data using boundary generationabstractWe present a method to reconstruct the three-dimensional shape of a moving instance of a known object category in video data. We exploit state-of-the-art semantic segmentation techniques to extract the object's two-dimensional shape in each frame. Therefore, our method is robust to occlusion, handles stationary objects and extends naturally to multiple video sequences. We apply Structure from Motion (SfM) to previously generated object images in order to compute a three-dimensional representation of the object. Our approach allows us to remove outliers in SfM reconstructions and to compute clean object meshes by leveraging previously computed semantic segmentations and virtual camera positions. We evaluate the accuracy of our method using a multi-view dataset of a moving vehicle. A laser scan serves as ground truth. We applied our algorithm on publicly available video data and on 25 sequences from our dataset. The algorithm achieves an average point distance of 3.3 cm evaluated on seven trajectories contained in the dataset. Sebastian Bullinger, Christoph Bodensteiner, Sebastian Wuttke, Michael Arens |
ICPR | 1 |