VLDB 2026 Research / reviewers in the wild / expert
Sebastian Krebs
dblp:61/8911
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-2740-6719ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Image recognition and object detection · 54% Video understanding and tracking · 22% Autonomous driving · 22% | |
| Human-computer interaction and pervasive computing
1 paper |
Immersive interaction · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection |
1.1 | 2 | 2024 | EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in Traffic · IEEE Trans. Pattern Anal. Mach. Intell. 2024 EuroCity Persons: A Novel Benchmark for Person Detection in Traffic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles |
0.8 | 1 | 2024 | EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in Traffic · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Video understanding and tracking › object tracking
person tracking |
0.8 | 1 | 2024 | EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in Traffic · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Image recognition and object detection › object detection › object detection evaluation
object detection benchmark |
0.4 | 1 | 2019 | EuroCity Persons: A Novel Benchmark for Person Detection in Traffic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Computer vision › Image recognition and object detection
pedestrian detection |
0.4 | 1 | 2019 | EuroCity Persons: A Novel Benchmark for Person Detection in Traffic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Immersive interaction › virtual reality
virtual reality storytelling |
0.2 | 1 | 2016 | SwiVRChair: A Motorized Swivel Chair to Nudge Users' Orientation for 360 Degree Storytelling in Virtual Reality · CHI 2016 |
Virtual and augmented reality
cybersickness |
0.1 | 1 | 2016 | SwiVRChair: A Motorized Swivel Chair to Nudge Users' Orientation for 360 Degree Storytelling in Virtual Reality · CHI 2016 |
Methods — techniques the papers use, named apart from their topics
semi-supervised pseudo ground-truth generation · 0.8ego-motion compensation · 0.8LiDAR-based 3D uplifting · 0.8user study · 0.5motorized swivel chair · 0.5haptic nudging · 0.5YOLOv3 · 0.4SSD · 0.4R-FCN · 0.4Faster R-CNN · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DenseBEV: Transforming BEV Grid Cells into 3D ObjectsabstractIn current research, Bird’s-Eye-View (BEV)-based transformers are increasingly utilized for multi-camera 3D object detection. Traditional models often employ random queries as anchors, optimizing them successively. Recent advancements complement or replace these random queries with detections from auxiliary networks. We propose a more intuitive and efficient approach by using BEV feature cells directly as anchors. This end-to-end approach leverages the dense grid of BEV queries, considering each cell as a potential object for the final detection task. As a result, we introduce a novel two-stage anchor generation method specifically designed for multi-camera 3D object detection. To address the scaling issues of attention with a large number of queries, we apply BEV-based Non-Maximum Suppression, allowing gradients to flow only through non-suppressed objects. This ensures efficient training without the need for post-processing. By using BEV features from encoders such as BEVFormer directly as object queries, temporal BEV information is inherently embedded. Building on the temporal BEV information already embedded in our object queries, we introduce a hybrid temporal modeling approach by integrating prior detections to further enhance detection performance. Evaluating our method on the nuScenes dataset shows consistent and significant improvements in NDS and mAP over the baseline, even with sparser BEV grids and therefore fewer initial anchors. It is particularly effective for small objects, enhancing pedestrian detection with a 3.8% mAP increase on nuScenes and an 8% increase in LET-mAP on Waymo. Applying our method, named DenseBEV, to the challenging Waymo Open dataset yields state-of-the-art performance, achieving a LET-mAP of 60.7%, surpassing the previous best by 5.4%. Code is available at https://github.com/mdaehl/DenseBEV. Marius Dähling, Sebastian Krebs, Johann Marius Zöllner |
WACV | 2 |
| 2025 | CrowdQuery: Density-Guided Query Module for Enhanced 2D and 3D Detection in Crowded ScenesabstractThis paper introduces a novel method for end-to-end crowd detection that leverages object density information to enhance existing transformer-based detectors. We present CrowdQuery (CQ), whose core component is our CQ module that predicts and subsequently embeds an object density map. The embedded density information is then systematically integrated into the decoder. Existing density map definitions typically depend on head positions or object-based spatial statistics. Our method extends these definitions to include individual bounding box dimensions. By incorporating density information into object queries, our method utilizes density-guided queries to improve detection in crowded scenes. CQ is universally applicable to both 2D and 3D detection without requiring additional data. Consequently, we are the first to design a method that effectively bridges 2D and 3D detection in crowded environments. We demonstrate the integration of CQ into both a general 2D and 3D transformer-based object detector, introducing the architectures CQ2D and CQ3D. CQ is not limited to the specific transformer models we selected. Experiments on the STCrowd dataset for both 2D and 3D domains show significant performance improvements compared to the base models, outperforming most state-of-the-art methods. When integrated into a state-of-the-art crowd detector, CQ can further improve performance on the challenging CrowdHuman dataset, demonstrating its generalizability. The code is released at https://github.com/mdaehl/CrowdQuery. Marius Dähling, Sebastian Krebs, Johann Marius Zöllner |
IROS | 2 |
| 2025 | Camera-and LiDAR-based Person Re-IdentificationabstractIn this paper, we introduce a novel method for creating appearance embeddings to identify individual persons using an object re-identification (ReID) framework. We present CLFormer (Camera LiDAR Transformer), a transformer-based architecture that incorporates multi-modal data from both camera and LiDAR sensors. We introduce the 3D Cuboid-Inclusive Point Embedding (3D-CIPE), which leverages rich data from LiDAR point clouds and 3D cuboids to add a learnable embedding into the transformer structure. Additionally, through ablation studies, we explore and analyze various strategies for the early and late fusion of multi-modal input data. To evaluate our proposed CLFormer, we reinterpret the nuScenes dataset [1] for ReID purposes and use it for our experiments. Our method demonstrates a significant improvement in performance, outperforming the image-only baseline with an increase of 2.3 in mean Average Precision (mAP). Sebastian Krebs, Dariu Gavrila |
IV | 1 |
| 2024 | EuroCity Persons 2.0: A Large and Diverse Dataset of Persons in TrafficabstractWe present the EuroCity Persons (ECP) 2.0 dataset, a novel image dataset for person detection, tracking and prediction in traffic. The dataset was collected on-board a vehicle driving through 29 cities in 11 European countries. It contains more than 250K unique person trajectories, in more than 2.0M images and comes with a size of 11 TB. ECP2.0 is about one order of magnitude larger than previous state-of-the-art person datasets in automotive context. It offers remarkable diversity in terms of geographical coverage, time of day, weather and seasons. We discuss the novel semi-supervised approach that was used to generate the temporally dense pseudo ground-truth (i.e., 2D bounding boxes, 3D person locations) from sparse, manual annotations at keyframes. Our approach leverages auxiliary LiDAR data for 3D uplifting and vehicle inertial sensing for ego-motion compensation. It incorporates keyframe information in a three-stage approach (tracklet generation, tracklet merging into tracks, track smoothing) for obtaining accurate person trajectories. We validate our pseudo ground-truth generation approach in ablation studies, and show that it significantly outperforms existing methods. Furthermore, we demonstrate its benefits for training and testing of state-of-the-art tracking methods. Our approach provides a speed-up factor of about 34 compared to frame-wise manual annotation. The ECP2.0 dataset is made freely available for non-commercial research use. Sebastian Krebs, Markus Braun 0003, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Simple Pair Pose - Pairwise Human Pose Estimation in Dense Urban Traffic ScenesabstractDespite the success of deep learning, human pose estimation remains a challenging problem in particular in dense urban traffic scenarios. Its robustness is important for followup tasks like trajectory prediction and gesture recognition. We are interested in human pose estimation in crowded scenes with overlapping pedestrians, in particular pairwise constellations. We propose a new top-down method that relies on pairwise detections as input and jointly estimates the two poses of such pairs in a single forward pass within a deep convolutional neural network. As availability of automotive datasets providing poses and a fair amount of crowded scenes is limited, we extend the EuroCity Persons dataset by additional images and pose annotations. With 46,975 images and poses of 279,329 persons our new EuroCity Persons Dense Pose dataset is the largest pose dataset recorded from a moving vehicle. In our experiments using this dataset we show improved performance for poses of pedestrian pairs in comparison with a state of the art method for human pose estimation in crowds. Markus Braun 0003, Fabian Flohr, Sebastian Krebs, Ulrich Kreße, Dariu Gavrila |
IV | 3 |
| 2020 | ECP2.5D - Person Localization in Traffic Scenesabstract3D localization of persons from a single image is a challenging problem, where advances are largely data-driven. In this paper, we enhance the recently released EuroCity Persons detection dataset, a large and diverse automotive dataset covering pedestrians and riders. Previously, only 2D annotations and image data were provided. We introduce an automatic 3D lifting procedure by using additional LiDAR distance measurements, to augment a large part of the reasonable subset of 2D box annotations with their corresponding 3D point positions (136K persons in 46K frames of day- and night-time). The resulting dataset (coined ECP2.5D), now including Li-DAR data as well as the generated annotations, is made publicly available for (non-commercial) benchmarking of camera-based and/or LiDAR 3D object detection methods. We provide baseline results for 3D localization from single images by extending the YOLOv3 2D object detector with a distance regression including uncertainty estimation. Markus Braun 0003, Sebastian Krebs, Dariu Gavrila |
IV | 2 |
| 2019 | EuroCity Persons: A Novel Benchmark for Person Detection in Traffic ScenesabstractBig data has had a great share in the success of deep learning in computer vision. Recent works suggest that there is significant further potential to increase object detection performance by utilizing even bigger datasets. In this paper, we introduce the EuroCity Persons dataset, which provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes. The images for this dataset were collected on-board a moving vehicle in 31 cities of 12 European countries. With over 238200 person instances manually labeled in over 47300 images, EuroCity Persons is nearly one order of magnitude larger than datasets used previously for person detection in traffic scenes. The dataset furthermore contains a large number of person orientation annotations (over 211200). We optimize four state-of-the-art deep learning approaches (Faster R-CNN, R-FCN, SSD and YOLOv3) to serve as baselines for the new object detection benchmark. We analyze the generalization capabilities of these detectors when trained with the new dataset. We furthermore study the effect of the training set size, the dataset diversity (day- vs. night-time, geographical region), the dataset detail (i.e. availability of object orientation information) and the annotation quality on the detector performance. Finally, we analyze error sources and discuss the road ahead. Markus Braun 0003, Sebastian Krebs, Fabian Flohr, Dariu Gavrila |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | SwiVRChair: A Motorized Swivel Chair to Nudge Users' Orientation for 360 Degree Storytelling in Virtual RealityabstractWe present SwiVRChair, a motorized swivel chair to nudge users' orientation in 360 degree storytelling scenarios. Since rotating a scene in virtual reality (VR) leads to simulator sickness, storytellers currently have no way of controlling users' attention. SwiVRChair allows creators of 360 degree VR movie content to be able to rotate or block users' movement to either show certain content or prevent users from seeing something. To enable this functionality, we modified a regular swivel chair using a 24V DC motor and an electromagnetic clutch. We developed two demo scenarios using both mechanisms (rotate and block) for the Samsung GearVR and conducted a user study (n=16) evaluating the presence, enjoyment and simulator sickness for participants using SwiVRChair compared to self control (Foot Control). Users rated the experience using SwiVRChair to be significantly more immersive and enjoyable whilst having a decrease in simulator sickness. Jan Gugenheimer, Dennis Wolf 0002, Gabriel Haas 0001, Sebastian Krebs, Enrico Rukzio |
CHI | 4 |
| 2014 | Probabilistic inference of visibility conditions by means of sensor fusionabstractWith the help of advanced driver assistance systems (ADAS), today's vehicles are already able to perform impressive perception tasks. Besides information about other traffic participants, the current environmental visibility condition is one key aspect to enable further development, especially in difficult scenarios and adverse weather conditions. This work presents a system to estimate the visibility range for both the driver and vision-based ADAS. On the basis of an existing probabilistic radar-camera vehicle tracking framework, individual visibility range measurements are deduced by monitoring camera measurements to vehicles already confirmed by the radar sensor. This individual track-level information is then combined with spatial and temporal memory to build a holistic system to infer the current visibility condition in a probabilistic way. Experiments on both synthetic and real-world data validate the proposed concepts. In addition, a conducted user study compares system outputs to human visibility perception on realword scenes. Michael Gabb, Sebastian Krebs, Otto Löhlein, Martin Fritzsche |
Intelligent Vehicles Symposium | 2 |